Image processing method and related device
By using semantic models and video segment vector indexing technology in electronic devices, the problem of numerous and complicated image files in the gallery is solved, and a faster and more accurate image search effect is achieved.
Patent Information
- Application Number
- CN202311585461.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2025-05-30
AI Technical Summary
With the increase in storage space of electronic devices, there are many and diverse images stored in the gallery, making it difficult for users to accurately find the desired image file.
Through semantic models, the user's search intention is matched with the image files in the gallery, and video segment vectors are generated and indexed to realize segmentation and semantic analysis of videos, so as to find image files faster and more accurately.
Through the analysis of semantic model, the image files that users want can be found faster and more accurately, improving the efficiency and accuracy of image search.
Smart Images

Figure CN120067356A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of terminals, and in particular, to an image processing method and related devices. Background Art
[0002] The gallery application in an electronic device can have an image management function and an image search function, and users can use the gallery to view or search for images. However, as the storage space of electronic devices becomes larger and larger, the number of images stored in the gallery is also increasing, and users cannot accurately find the images they want. Summary of the Invention
[0003] The image processing method and related devices provided in the embodiments of this application enable an electronic device to match the image files that a user wants to find with the image files in the gallery based on a semantic model, so as to provide the user with images that are closer to the search content. In this way, by analyzing the user's search intention through the semantic model, the image files that the user wants can be found faster and more accurately.
[0004] In a first aspect, an image processing method provided in the embodiments of this application includes:
[0005] An electronic device displays a first interface of a first application, and the first interface includes an input box; in response to an operation in which a user inputs text in the input box, the electronic device displays a second interface, and the second interface includes one or more thumbnails; among them, the thumbnails include thumbnails of videos, and the thumbnails of videos are obtained by matching from an index library based on vectors of the input text. The index library stores video segment vectors corresponding to videos in the first application, and the video segment vectors corresponding to videos in the first application are obtained in advance in the following manner: the videos in the first application are segmented for the first time according to a preset duration to obtain video segments after the first segmentation; the video segments after the first segmentation are segmented for the second time according to the image similarity of adjacent frames to obtain video segments after the second segmentation; video segment vectors are generated for the video segments after the second segmentation. In this way, the image files that the user wants can be found faster and more accurately through vector similarity analysis.
[0006] In a possible implementation, before performing a second video segmentation on the video segments after the first segmentation according to the image similarity of adjacent frames, it further includes: decoding the video segments after the first segmentation to obtain the decoded video segments; extracting one or more video frames from the decoded video segments; performing a second video segmentation on the video segments after the first segmentation according to the image similarity of adjacent frames, including: performing video frame analysis on one or more video frames to obtain the image similarity of adjacent video frames; if the image similarity of adjacent video frames meets the similarity threshold, dividing the adjacent video frames into the same video segment, and if the image similarity of adjacent video frames does not meet the similarity threshold, dividing the adjacent video frames into different video segments. In this way, the video can be segmented, which is convenient for the semantic model to analyze the video more accurately and convenient for users to search.
[0007] In a possible implementation, extracting one or more video frames from the decoded video segments includes: extracting one or more video frames from the decoded video segments according to the frame extraction strategy, and the frame extraction strategy includes: extracting one or more video frames from the decoded video segments according to key frames, or extracting one or more video frames from the decoded video segments according to consecutive frames. In this way, the video frames can include image information, can represent the semantics of the video, and are convenient for subsequent analysis of the video frames, and thus more accurately implement video semantic search.
[0008] In a possible implementation, the video frame analysis includes one or more of the following: label detection, clarity detection, or jitter detection. In this way, based on the algorithm operation strategy for video frame analysis, the calculation method of video frame analysis can be flexibly adjusted, so that while considering the calculation speed, the power consumption of the electronic device and the occurrence of lags can be reduced as much as possible.
[0009] In a possible implementation, after responding to the user's operation of inputting text in the input box, it further includes: generating a text vector for the text; the electronic device displays a second interface, and the second interface includes one or more thumbnails, including: judging the similarity between the text vector and the video segment vector; if the similarity between the text vector and the video segment vector is greater than or equal to the first threshold, the electronic device displays a second interface, and the second interface includes thumbnails of the video. In this way, the electronic device can match the image file that the user wants to search with the image files in the image library based on the semantic vector, so as to provide the user with an image closer to the search content. The user can find the desired image file faster and more accurately.
[0010] In a possible implementation, before generating video segment vectors for the video segments after the second segmentation, it further includes: saving the relevant information of the video into an index library, where the relevant information of the video includes one or more of the following: the hash value of the video, the path of the video, or the first timestamp of the video; after generating video segment vectors for the video segments after the second segmentation, it further includes: sending a first instruction to the file library storing the videos in the first application; the file library updates the timestamp of the video to a second timestamp based on the first instruction; it is determined that the second timestamp is later than the first timestamp, and the relevant information of the video is updated in the index library, where the relevant information of the video includes one or more of the following: the hash value of the video, the video segment vector, the path of the video, or the third timestamp of the video, and the third timestamp is later than the second timestamp. In this way, by timely updating the timestamp information in the media index library and the media file library, the analysis status of the current video can be accurately judged, and the process of generating semantic vectors and indexes for a certain video will not be repeated, so that the image file desired by the user can be found faster and more accurately.
[0011] In a possible implementation, before generating video segment vectors for the video segments after the second segmentation, it further includes: saving the relevant information of the video into an index library, where the relevant information of the video includes one or more of the following: the hash value of the video, the path of the video, or the first timestamp of the video; after generating video segment vectors for the video segments after the second segmentation, it further includes: sending a first instruction to the index library; the index library updates the relevant information of the video based on the first instruction, where the relevant information of the video includes one or more of the following: the hash value of the video, the video segment vector, the path of the video, or the second timestamp of the video, and the second timestamp is later than the first timestamp. In this way, the image file desired by the user can be found faster and more accurately.
[0012] In a possible implementation, the preset duration is obtained from the cloud server or determined based on the resources of the codec. In this way, by obtaining the preset duration from the cloud server, the cloud server can flexibly modify the value of the second preset duration according to the actual situation for different electronic devices, improving the flexibility and compatibility of the solution. Determining the preset duration based on the resources of the codec can reduce the resource occupancy of each video slice on the codec, so that the codec is not prone to jamming when decoding the video.
[0013] Second aspect, an embodiment of the present application provides an image processing device, which may be an electronic device, or a chip or a chip system inside the electronic device. The device may include a processing unit and a display unit. The processing unit is used to implement any processing-related method executed by the electronic device in the first aspect or any possible implementation manner of the first aspect. The display unit is used to implement any display-related method executed by the electronic device in the first aspect or any possible implementation manner of the first aspect. When the device is an electronic device, the processing unit may be a processor. The device may further include a storage unit, and the storage unit may be a memory. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to enable the electronic device to implement the method described in the first aspect or any possible implementation manner of the first aspect. When the device is a chip or a chip system inside the electronic device, the processing unit may be a processor. The processing unit executes the instructions stored in the storage unit to enable the electronic device to implement the method described in the first aspect or any possible implementation manner of the first aspect. The storage unit may be a storage unit inside the chip (such as registers, caches, etc.), or a storage unit outside the chip and inside the electronic device (such as read-only memory, random access memory, etc.).
[0014] Exemplarily, the display unit is used to display the first interface of the first application and is further used to display the second interface. The processing unit is used to perform the first video segmentation on the video in the first application according to a preset duration; is further used to perform the second video segmentation on the video segments after the first segmentation according to the image similarity of adjacent frames, and is further used to generate video segment vectors for the video segments after the second segmentation respectively.
[0015] In a possible implementation manner, the processing unit is used to decode the video segments after the first segmentation, is further used to extract one or more video frames from the decoded video segments, and is further used to perform video frame analysis on the one or more video frames. Specifically, it is further used to divide adjacent video frames into the same video segment or divide adjacent video frames into different video segments.
[0016] In a possible implementation manner, the processing unit is used to extract one or more video frames from the decoded video segments according to a frame extraction strategy.
[0017] In a possible implementation manner, the video frame analysis includes one or more of the following: label detection, clarity detection, or jitter detection.
[0018] In a possible implementation manner, the processing unit is used to generate a text vector for the text and is further used to determine the similarity between the text vector and the video segment vector. The display unit is used to display the second interface, and the second interface includes a thumbnail of the video.
[0019] In a possible implementation, a processing unit is configured to save relevant information of a video into an index library, and is further configured to send a first instruction to a file library storing videos in a first application; and is further configured to update the timestamp of the video to a second timestamp; specifically, it is further configured to update the relevant information of the video in the index library.
[0020] In a possible implementation, a processing unit is configured to save relevant information of a video into an index library, and is further configured to send a first instruction to the index library; specifically, it is further configured to update the relevant information of the video.
[0021] In a possible implementation, the preset duration is obtained from a cloud-side server or determined based on the resources of a codec.
[0022] In a third aspect, an embodiment of the present application provides a terminal device, including a processor and a memory. The memory is used to store code instructions, and the processor is used to run the code instructions to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0023] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. A computer program or instruction is stored in the computer-readable storage medium. When the computer program or instruction runs on a computer, the computer is caused to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0024] In a fifth aspect, an embodiment of the present application provides a computer program product including a computer program. When the computer program runs on a computer, the computer is caused to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0025] In a sixth aspect, the present application provides a chip or a chip system. The chip or the chip system includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected by a line. The at least one processor is used to run a computer program or instruction to execute the method described in the first aspect or any possible implementation manner of the first aspect. Among them, the communication interface in the chip may be an input / output interface, a pin, a circuit, etc.
[0026] In a possible implementation, the chip or the chip system described above in the present application further includes at least one memory, and instructions are stored in the at least one memory. The memory may be an internal storage unit of the chip, for example, a register, a cache, etc., or may be a storage unit of the chip (for example, a read-only memory, a random access memory, etc.).
[0027] It should be understood that the second aspect to the sixth aspect of the present application correspond to the technical solutions of the first aspect of the present application, and the beneficial effects obtained by each aspect and the corresponding feasible implementation manners are similar and will not be repeated. Brief Description of the Drawings
[0028] Figure 1 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0029] Figure 2 FIG. is a schematic software structure diagram of an electronic device provided by an embodiment of the present application;
[0030] Figure 3 FIG. is a schematic diagram of a search interface of a picture gallery provided by an embodiment of the present application;
[0031] Figure 4 FIG. is a timing diagram of generating a semantic vector of a picture provided by an embodiment of the present application;
[0032] Figure 5 FIG. is a schematic diagram of a picture gallery interface provided by an embodiment of the present application;
[0033] Figure 6 FIG. is a schematic diagram of a progress bar interface of a picture gallery provided by an embodiment of the present application;
[0034] Figure 7 FIG. is a timing diagram of generating a semantic vector of a video segment provided by an embodiment of the present application;
[0035] Figure 8 FIG. is a schematic diagram of video slicing and video segmentation provided by an embodiment of the present application;
[0036] Figure 9 FIG. is a timing diagram of index construction provided by an embodiment of the present application;
[0037] Figure 10 FIG. is a timing diagram of user retrieval provided by an embodiment of the present application;
[0038] Figure 11 FIG. is a schematic diagram of a media semantic search framework provided by an embodiment of the present application;
[0039] Figure 12 FIG. is a schematic diagram of an image processing method provided by an embodiment of the present application;
[0040] Figure 13 FIG. is a schematic structural diagram of a chip provided by an embodiment of the present application. Detailed Description of the Embodiments
[0041] For the convenience of clearly describing the technical solutions of the embodiments of the present application, the following briefly introduces some terms and technologies involved in the embodiments of the present application:
[0042] 1. The semantic model can also be referred to as a large semantic model or a semantic algorithm model, and this semantic model can be deployed on the NPU chip platform. The semantic model can provide a contrastive language-image pre-training (CLIP) model. Based on the text encoder and image encoder of the CLIP model, the semantic model conducts contrastive learning to train a text encoder for outputting text semantic vectors of text and an image encoder for outputting image semantic vectors of pictures or video frames.
[0043] In a possible implementation, the semantic model can output the picture semantic vector of the picture through the image encoder based on the picture information. The picture semantic vector can be understood as the eigenvalue of the picture, and the eigenvalue of the picture can be used to represent the semantics of the picture. Thus, the picture semantic vector can be used to identify the picture.
[0044] 2. Terms
[0045] In the embodiments of this application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. For example, the first chip and the second chip are only used to distinguish different chips and do not limit their sequence. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different.
[0046] It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way.
[0047] In the embodiments of this application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the front and rear associated objects. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.
[0048] 3. Electronic device
[0049] The electronic device according to the embodiments of the present application may also be any form of terminal device. For example, the electronic device may include: mobile phone, tablet computer, handheld computer, laptop computer, mobile internet device (MID), wearable device, virtual reality (VR) device, augmented reality (AR) device, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, wireless terminal in smart home, cellular phone, cordless phone, session initiation protocol (SIP) phone, wireless local loop (WLL) station, personal digital assistant (PDA), handheld device with wireless communication function, computing device or other processing device connected to a wireless modem, vehicle-mounted device, wearable device, electronic device in a 5G network or electronic device in a future evolved public land mobile network (PLMN), etc. The embodiments of the present application are not limited thereto.
[0050] By way of example and not limitation, in the embodiments of the present application, the electronic device may also be a wearable device. A wearable device may also be referred to as a wearable intelligent device, which is a general term for devices developed by applying wearable technologies to intelligentize daily wear, such as glasses, gloves, watches, clothing, and shoes. A wearable device is a portable device that is either directly worn on the body or integrated into the user's clothes or accessories. A wearable device is not only a hardware device, but also realizes powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable intelligent devices include those with complete functions and large sizes that can achieve complete or partial functions without relying on a smart phone, such as smart watches or smart glasses, etc., and those that only focus on a certain type of application function and need to cooperate with other devices such as smart phones, such as various smart bracelets and smart jewelry for monitoring physical signs.
[0051] In addition, in the embodiments of the present application, the electronic device may also be an electronic device in an Internet of Things (IoT) system. The IoT is an important part of the future development of information technology. Its main technical feature is to connect objects to the network through communication technology, thereby realizing an intelligent network of human-machine interconnection and object-object interconnection.
[0052] The electronic device in the embodiments of the present application may also be referred to as: user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile platform, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent, or user device, etc.
[0053] In the embodiments of the present application, the electronic device or each network device includes a hardware layer, an operating system layer running on the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and a memory (also referred to as main memory). The operating system can be any one or more computer operating systems that implement service processing through processes. For example, Linux operating system, Unix operating system, Android operating system, iOS operating system, or Windows operating system, etc. The application layer includes applications such as a browser, an address book, a word processing software, and an instant messaging software.
[0054] Exemplarily, Figure 1 shows a schematic structural diagram of the electronic device.
[0055] The electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0056] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown, or combine certain components, or split certain components, or have different component arrangements. The components shown may be implemented by hardware, software, or a combination of software and hardware.
[0057] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, an accelerated processing unit (APU), and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. The controller may generate operation control signals according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0058] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can be directly called from the above-mentioned memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0059] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are only illustrative descriptions and do not constitute a structural limitation on the electronic device. In other embodiments of the present application, the electronic device may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0060] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc. The data storage area can store the data created during the use of the electronic device. The processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor. For example, in the embodiments of the present application, the internal memory 121 can store instructions for executing an image processing method. The processor 110 can implement the processing of images by executing the instructions stored in the internal memory 121.
[0061] The display screen 194 is used to display photos, videos, etc. The display screen 194 includes a display panel. In some embodiments, the electronic device may include one or N display screens 194, where N is a positive integer greater than 1. The electronic device realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, and is connected to the display screen 194 and the application processor. For example, in the embodiments of the present application, the gallery application can display photos and videos through the display screen 194.
[0062] The GPU is used to execute mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information. The electronic device realizes the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc. For example, in the embodiments of the present application, the pictures and videos stored in the gallery application can be taken and stored by the user using the electronic device.
[0063] The NPU can be used to implement the intelligent cognition of electronic devices, such as image recognition, face recognition, speech recognition, text understanding, etc. For example, in the embodiments of the present application, the NPU can perform text understanding on the search text input by the user in the search interface of the gallery.
[0064] Figure 2 It is the software structure block diagram of the electronic device in the embodiments of the present application. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom are the application layer, the application framework layer, Android runtime and system libraries, the hardware adaptation layer (HAL), and the kernel layer.
[0065] The application layer can also be called the app layer, and the app layer can include a series of application packages. As Figure 2 shown, the application packages can include applications such as the gallery, video, camera, etc. In some scenarios, the gallery can also be called the photo album. The applications can include system applications and third-party applications.
[0066] Among them, the app layer can include the business layer and the application function layer.
[0067] For example, in the embodiments of the present application, the business layer can be used to execute business logics related to the interaction between the gallery and the user. The user can input text information in the input box of the gallery, and the business layer can also display the results of semantic image search for the user on the gallery interface, etc.
[0068] The application function layer can run computer vision (CV) algorithm services. In some implementations, this CV algorithm service can also be called a CV algorithm pipeline, a CV algorithm service pipeline, or a media file analysis service. The CV algorithm service can include processing nodes such as label analysis, portrait analysis, and image scoring.
[0069] Label analysis can be used to classify tags for pictures or videos. For example, the tag types can include people, animals, landscapes, food, etc. Label analysis can roughly classify pictures or videos.
[0070] Portrait analysis can be used to extract face features and perform face detection, and can also identify information such as faces and ages according to the large model, so as to distinguish different people. In possible scenarios, the gallery can create a photo album for a certain person alone based on portrait analysis.
[0071] Image scoring can generate an image score from aspects such as image quality and aesthetics based on a quality assessment algorithm. The image score can evaluate the clarity of the image, the proportion of people in the image, etc. For example, the larger the value of the image score, the clearer the image, and the smaller the value of the image score, the blurrier the image. In a possible scenario, the image library can filter out photos with higher quality based on the image score, so that these higher-quality images can be used as the cover of the photo album, or these higher-quality images can also be used.
[0072] In the embodiment of the present application, the CV algorithm service may further include processing nodes such as semantic analysis and video slicing.
[0073] Semantic analysis can be used to perform image feature analysis on pictures or video segments, so as to generate semantic vectors corresponding to the pictures or video segments.
[0074] Video slicing can be used to segment a longer video.
[0075] The application function layer may further include a media file library, a semantic vector library, and a video segment information library. Among them, the media file library can store the images in the image library. The semantic vector library can store the semantic vectors of the pictures after semantic analysis, and the semantic vector library can also be called a media vector library. The video segment information library can store video segment information. The application program framework layer can also be called the Framework layer. The Framework layer can provide application programming interfaces (APIs) and programming frameworks for the application programs in the application layer. The Framework layer can include some predefined functions.
[0076] As Figure 2 shown, the Framework layer can include a smart algorithm middle platform, a window manager, a resource manager, a content provider, and a view system, etc.
[0077] Among them, the smart algorithm middle platform can be used to provide image algorithm analysis services for the image library. The smart algorithm middle platform can include a smart algorithm module, an index construction module, etc.
[0078] The smart algorithm module can include a semantic algorithm module, etc. The semantic algorithm module can perform semantic analysis on pictures or videos, output picture semantic vectors or video semantic vectors, and store the semantic vectors in the semantic vector library.
[0079] The index construction module can generate corresponding indexes for pictures or videos and store the indexes in the media index library, which is convenient for subsequently finding pictures or videos with high similarity in the media index library. The index construction module may include a natural language understanding (NLU) semantic extraction module, which can perform semantic extraction on the information input by the user.
[0080] The Android runtime includes a core library and a virtual machine. The Android runtime is responsible for the control and management of the Android system.
[0081] The core library consists of two parts: one is the functional functions that need to be called by the Java language, and the other is the core library of Android.
[0082] The application layer and the Framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the Framework layer as binary files. The virtual machine is used to perform functions such as management of object life cycles, stack management, thread management, security and exception management, and garbage collection. For example, in the embodiments of the present application, the virtual machine can be used to perform functions such as image semantic analysis, video segmentation, and construction of vector indexes.
[0083] The system library can also be referred to as the Native layer, and the Native layer can include multiple functional modules. For example: media library, function library, graphics processing library, etc.
[0084] The kernel layer is the layer between hardware and software. The kernel layer can include NPU drivers, CPU drivers, GPU drivers, and / or APU drivers, etc.
[0085] It should be noted that the embodiments of the present application only take the Android system as an example for illustration. In other operating systems (such as Windows system, IOS system, etc.), as long as the functions implemented by each functional module are similar to those of the embodiments of the present application, the solution of the present application can also be implemented.
[0086] The electronic device may include a gallery application, which can have image management functions and image search functions. Users can use the gallery to view or search for images, where the images can include pictures and videos. As the storage space of the electronic device becomes larger and larger, the number of images stored in the gallery is also increasing, and the image search function is used more and more frequently.
[0087] In some implementations, the image file may have attribute tags such as shooting time, shooting location, image name, etc. When the user enters search text such as time, location, etc. in the search box of the gallery for searching, the electronic device can display the image files that match the search text entered by the user.
[0088] However, the search method based on attribute tags is relatively simple. It can only roughly search and classify image files and cannot analyze the specific content contained in the image files. When the user inputs complex search text, the electronic device may not be able to display the image files that exactly correspond to the user's needs, resulting in inaccurate found image files.
[0089] In view of this, in the image processing method provided by the embodiments of the present application, the electronic device can match the image files that the user wants to find with the image files in the image library based on a semantic model, so as to provide the user with images that are closer to the search content. In this way, by analyzing the user's search intention through the semantic model, the image files that the user wants can be found faster and more accurately.
[0090] Exemplarily, Figure 3 FIG. 301 is a search interface of the image library. The search interface 301 may include a search bar 302 and a search button 303. The user can input search text in the search bar 302. For example, the search text can be "a cat staring intently".
[0091] After the user inputs the search text, the search button 303 can be triggered. In response to the user's trigger operation, the electronic device can perform semantic analysis according to the user's search text, convert the search text into a text semantic vector, and perform similarity matching based on the text semantic vector and the semantic vectors corresponding to the pictures or videos. The electronic device can use the pictures and / or videos with a similarity greater than or equal to a certain preset threshold as search results and display the thumbnails of the pictures and / or videos on the search interface 301.
[0092] Among them, the search results in the interface 301 may include a first display area 304, a second display area 305, and a third display area 306. In the first display area 304, the electronic device can display the picture search results that meet the similarity requirements and display the number "569" of the picture search results. In the second display area 305, the electronic device can display the photo search results that meet the similarity requirements and contain text, and display the number "99" of the photo search results that contain text. In the third display area 306, the electronic device can display the video search results that meet the similarity requirements and display the number "88" of the video search results. The matching video frames can also be displayed as thumbnails. The third display area 306 can also display the time point of the video. The time point of the video can be the time point of the video segment to which the video frame corresponding to the thumbnail belongs. For example, the time point can be the time point of the video frame corresponding to the thumbnail, or the time point of other video frames in the video segment. The embodiments of the present application do not make any limitations.
[0093] Exemplarily, taking the search text as "a cat staring intently" as an example, the found image files may include: pictures with relevant content of "a cat staring intently" in the picture, pictures with relevant pictures whose text in the picture contains keywords such as "staring intently" and "cat", videos with relevant content of "a cat staring intently" in a certain frame picture of the video, and / or videos with relevant pictures in a certain frame picture of the video whose text contains keywords such as "staring intently" and "cat", etc.
[0094] It can be understood that if the image contains text, the electronic device can identify the text content in the image based on the character recognition technology, so as to determine whether the text in the image contains the keywords of the search text. If the text in the image contains the keywords of the search text, the image meets the search requirements, otherwise it means that the image does not meet the search requirements.
[0095] The image processing method provided by the embodiments of the present application may include (1) generating a semantic vector of the image, (2) constructing an index of the semantic vector, and (3) retrieving text semantics.
[0096] (1) Generate a semantic vector of the image.
[0097] The image processing method of the embodiments of the present application can generate both picture semantic vectors and video semantic vectors. Among them, the relevant process of generating picture semantic vectors can refer to the following Figure 4 relevant description, and the relevant process of generating video semantic vectors can refer to the following Figure 7 relevant description.
[0098] It can be understood that the execution order of generating picture semantic vectors and generating video semantic vectors can be not limited. It can generate semantic vectors for pictures first, or generate semantic vectors for videos first, or generate semantic vectors for pictures and videos simultaneously. In a possible implementation, since most of the images in the image library may be pictures, that is, the number of pictures is more than the number of videos, therefore, semantic vectors can be generated for pictures first, and then semantic vectors can be generated for videos. In this way, the image library can quickly respond to the user's search operation and improve the user experience.
[0099] Exemplarily, Figure 4 shows a timing diagram of generating a semantic vector of a picture.
[0100] S401. The user triggers a charging and / or screen-off operation.
[0101] S402. The image library monitors the state of the electronic device.
[0102] S403. The image library judges the temperature and / or power of the electronic device.
[0103] It can be understood that the picture gallery can monitor some states of the electronic device. For example, the states can include the charging state, the screen-off state, etc.
[0104] After the user triggers the charging and / or screen-off operation, the picture gallery can obtain the charging state and / or screen-off state of the electronic device, and then can trigger the start of the CV algorithm service for CV algorithm analysis.
[0105] It can be understood that, in order to avoid the electronic device having too high a temperature or too low a battery power, the picture gallery can start the CV algorithm service when it determines that the temperature of the electronic device is less than the temperature threshold and / or the battery power of the electronic device is less than the power threshold. In this way, the temperature of the electronic device can be prevented from being too high, and the battery life of the electronic device can be maintained, thereby improving the user experience.
[0106] Optionally, in addition to detecting the charging state and / or screen-off state of the electronic device, the picture gallery can also start the CV algorithm service for CV algorithm analysis according to the active trigger of the user.
[0107] Exemplarily, the electronic device can display the picture gallery interface. When there are images in the picture gallery that need to be analyzed by the CV algorithm, the picture gallery interface can be as Figure 5 shown in interface 501. Interface 501 can include a first prompt 502, a second prompt 503, an immediate start button 504, and multiple albums, etc.
[0108] Among them, the content of the first prompt 502 can include "Intelligent Image Recognition, Multiple Wonders", and the content of the second prompt 503 can include "Can Generate Memories, Intelligent Clustering of Photos, Making Search More Precise". Specifically, the content of the first prompt 502 and the content of the second prompt 503 are not limited in the embodiments of the present application.
[0109] The multiple albums can include a "Camera" album, an "All Photos" album, a "Videos" album, a "Screenshots & Screen Recordings" album, a "My Favorites" album, a "Recovery" album, and a "Recently Deleted" album, etc.
[0110] When the user clicks the immediate start button 504, the electronic device can respond to the user's click operation, start to recognize the image, and display an interface such as Figure 6 interface 601. Interface 601 can include a progress bar 602 for image recognition, a pause button 603, a prompt message 604, etc. Among them, the content of the prompt message 604 can include "Estimated remaining XX minutes" or "Intelligent Image Recognition in Progress, It is recommended to connect to power supply", etc. Specifically, the content of the prompt message 604 is not limited in the embodiments of the present application.
[0111] It can be understood that in the interface 501, if the user does not want to perform image recognition, the user can perform a preset operation. For example, the preset operation may include operations such as swiping up or down the interface. The specific preset operation can be pre-set by the photo gallery, and the embodiments of the present application do not make any limitations. In response to the user's preset operation, the electronic device can cancel the display of the first prompt 502, the second prompt 503, and the start immediately button 504 in the interface 501.
[0112] Alternatively, the interface 501 may further include a cancel button. When the user clicks the cancel button, in response to the user's click operation, the electronic device can cancel the display of the first prompt 502, the second prompt 503, and the start immediately button 504 in the interface 501.
[0113] In the interface 601, if the user does not want to perform image recognition, the user can click the pause button 603. In response to the user's click operation, the electronic device can stop recognizing the image in the interface 601.
[0114] It can be understood that the electronic device can also trigger CV algorithm analysis in other scenarios. For example, the electronic device can also trigger CV algorithm analysis when it is in a charging state and the user has not operated the electronic device for a long time. The specific scenarios for triggering CV algorithm analysis are not limited in the embodiments of the present application.
[0115] S404. The photo gallery calls the interface of the service interface layer to start the CV process.
[0116] S405. The service interface layer creates a CV pipeline thread.
[0117] S406. The CV pipeline registers algorithm nodes.
[0118] In the CV pipeline thread, algorithm nodes can be registered. Among them, the algorithm nodes can include semantic analysis nodes, label analysis nodes, portrait analysis nodes, image scoring nodes, video slicing nodes, etc.
[0119] The CV pipeline thread can create threads for each node and execute the algorithm processes corresponding to each node. For example, the CV pipeline thread can create a semantic analysis node thread and run the semantic analysis algorithm. The CV pipeline thread can create a label analysis node thread and run the label analysis algorithm. The CV pipeline thread can create a portrait analysis node thread and run the portrait analysis algorithm. The CV pipeline thread can create an image scoring node thread and run the image scoring algorithm. The CV pipeline thread can create a video slicing node thread and run the video slicing algorithm. Among them, the embodiments of the present application do not limit the execution order of each node.
[0120] For the convenience of description, the following takes the CV pipeline thread creating a semantic analysis node thread and running the semantic analysis algorithm as an example for illustration.
[0121] S407. The CV pipeline creates a semantic analysis node thread and runs the semantic analysis algorithm.
[0122] The CV pipeline can create a semantic analysis node thread for an image and run the semantic analysis algorithm in the semantic analysis node thread. For details, refer to the relevant descriptions in steps S408 - S419 below.
[0123] S408. The CV pipeline queries the algorithm capabilities.
[0124] S409. The algorithm layer returns the algorithm capability status.
[0125] The CV pipeline can call the relevant interfaces of the algorithm adaptation layer to query the algorithm capabilities of semantic analysis. The algorithm adaptation layer can query the algorithm capabilities of semantic analysis in the algorithm layer through interaction with the algorithm interface layer. The algorithm layer can return the algorithm capability status to the algorithm interface layer. The algorithm adaptation layer can obtain the algorithm capability status through interaction with the algorithm interface layer and return the algorithm capability status to the CV pipeline.
[0126] Among them, the algorithm capability status can be used to indicate whether the electronic device supports the algorithm capabilities of semantic analysis. The algorithm capability status can be any possible data type such as integer, boolean, string, etc. For example, when the algorithm capability status is true, it can indicate that the electronic device supports the algorithm capabilities of semantic analysis; when the algorithm capability status is false, it can indicate that the electronic device does not support the algorithm capabilities of semantic analysis. The specific values of the algorithm capability status are not limited in the embodiments of this application.
[0127] S410. The CV pipeline determines the algorithm capability status.
[0128] The CV pipeline can determine whether it supports the algorithm capabilities of semantic analysis according to the algorithm capability status. The specific determination method of the algorithm capability status can refer to the relevant descriptions in step S409 above and will not be elaborated here.
[0129] If the CV pipeline determines that the electronic device does not support the algorithm capabilities of semantic analysis, it will not continue to execute the relevant processes of the semantic analysis node.
[0130] If the CV pipeline determines that the electronic device supports the algorithm capabilities of semantic analysis, it can further initialize the semantic model and load the image data. It should be noted that the initialization of the semantic model and the loading of the image data do not need to distinguish the execution order. The CV pipeline can initialize the semantic model first, or load the image data first, or the processes of initializing the semantic model and loading the image data can be executed in parallel. The embodiments of this application do not make any limitations.
[0131] For the convenience of description, the following takes the process of loading the semantic model first as an example for illustration.
[0132] S411. Load the semantic model.
[0133] S412. Return the semantic model loading status.
[0134] The CV pipeline can call the relevant interfaces of the algorithm adaptation layer to load the semantic model. Through interactions with the algorithm interface layer, algorithm layer, etc., the algorithm adaptation layer can load the semantic model of the chip layer. After the semantic model is loaded, the model loading status can be returned to the algorithm layer. Through interactions with the algorithm interface layer, algorithm layer, etc., the algorithm adaptation layer can obtain the model loading status and return the model loading status to the CV pipeline.
[0135] Among them, the model loading status can be used to indicate whether the semantic model is successfully loaded. The model loading status can be any possible data type such as integer, boolean, string, etc. For example, when the model loading status is true, it can indicate that the semantic model is successfully loaded; when the model loading status is false, it can indicate that the semantic model is not successfully loaded. The specific values of the model loading status are not limited in the embodiments of the present application.
[0136] In a possible implementation, while loading the semantic model, the CV pipeline can also execute steps S413 - S415 to load picture data from the media file library.
[0137] S413. The CV pipeline loads picture data from the media file library.
[0138] S414. The CV pipeline loads the first quantity of pictures from the media file library each time.
[0139] S415. The media file library returns the picture data.
[0140] The CV pipeline can load the first quantity of pictures from the media file library each time. In a possible implementation, the first quantity can be 2000. The first quantity can be preset by the CV pipeline. For example, the first quantity can be an experimental test value. The embodiments of the present application do not limit the specific value of the first quantity. Among them, the value of the first quantity should be able to satisfy the timeliness of semantic search and not affect the smooth operation of the electronic device, etc.
[0141] It can be understood that when the CV pipeline loads the first quantity of pictures from the media file library each time, it can first query whether the identification information of these pictures exists in the semantic vector library. For example, the identification information of the pictures can include the hash value of the pictures, etc.
[0142] If the identification information of a certain icon does not exist in the semantic vector library, it indicates that the semantic vector of the picture has not been generated, and the picture can be loaded for semantic analysis; if the identification information of a certain icon exists in the semantic vector library, it indicates that the semantic vector of the picture has been generated, and there is no need to load the picture. In this way, the pictures in the media file library will not be loaded repeatedly, nor will the semantic vectors be generated repeatedly.
[0143] When the CV pipeline loads pictures from the media file library, if the number of pictures that have not undergone semantic analysis does not meet the first quantity, the CV pipeline can either load only the pictures that have not undergone semantic analysis, or load both the pictures that have not undergone semantic analysis and the videos that have not undergone semantic analysis. The embodiments of the present application do not make any limitations in this regard.
[0144] S416. The CV pipeline determines the model loading status.
[0145] The CV pipeline can determine whether the model is successfully loaded based on the model loading status. The specific method for determining the model loading status can refer to the relevant description in step S412 above and will not be elaborated here.
[0146] If the CV pipeline determines that the semantic model has not been successfully loaded, it will not continue to execute the relevant processes of the semantic analysis node.
[0147] If the CV pipeline determines that the semantic model is successfully loaded and the picture data is successfully read, it can execute the semantic analysis algorithm on the picture.
[0148] S417. The CV pipeline calls the semantic analysis algorithm.
[0149] After the CV pipeline reads the picture data from the media file library, it can pass the picture data to the semantic model through the relevant interfaces of the algorithm adaptation layer, algorithm interface layer, algorithm layer, etc. The semantic model can execute the semantic analysis algorithm to perform semantic analysis on the picture.
[0150] S418. The semantic model returns the picture semantic vector.
[0151] The semantic model can generate the picture semantic vector and return the picture semantic vector to the CV pipeline through the relevant interfaces of the algorithm adaptation layer, algorithm interface layer, algorithm layer, etc.
[0152] S419. Save the picture semantic vector.
[0153] The CV pipeline can save the obtained picture semantic vector to the semantic vector library.
[0154] In the semantic vector library, the picture semantic vector can be saved in any possible way. In a possible implementation, the picture semantic vector in the semantic vector library can be saved as a 768-bit array. The specific method for saving the picture semantic vector in the semantic vector library is not limited in the embodiments of the present application.
[0155] It can be understood that the execution processes of the tag analysis node, the portrait analysis node, and the image scoring node are similar to that of the semantic analysis node, and will not be elaborated here. Executing the tag analysis node can return the tag information of the picture or video. For example, the tag information can include tag type information such as people, animals, landscapes, and delicacies. Executing the portrait analysis node can return face detection information. For example, the face detection information can include information such as faces and ages. Executing the image scoring node can return the scoring information of the picture or video generated from aspects such as image quality and aesthetics, etc.
[0156] Exemplarily, Figure 7 A timing diagram showing the generation of the semantic vector of the video segment is shown.
[0157] S701. The user triggers a charging and / or screen-off operation.
[0158] S702. The gallery monitors the state of the electronic device.
[0159] S703. The gallery judges the temperature and / or power of the electronic device.
[0160] S704. The gallery calls the interface of the service interface layer to start the CV process.
[0161] S705. The service interface layer creates a CV pipeline thread.
[0162] S706. Register the algorithm node.
[0163] It can be understood that steps S701 - S706 can refer to the relevant descriptions of steps S401 - S406 in the corresponding embodiment above, and will not be elaborated here. Figure 4 Corresponding to the relevant descriptions of steps S401 - S406 in the corresponding embodiment, and will not be elaborated here.
[0164] S707. The CV pipeline creates a semantic analysis node thread and runs the semantic analysis algorithm.
[0165] The CV pipeline can create a semantic analysis node thread for the video and run processes such as the video analysis algorithm and the semantic analysis algorithm in the semantic analysis node thread. Specifically, it can refer to the relevant descriptions of the following steps S708 - S427.
[0166] S708. The CV pipeline queries the algorithm capabilities.
[0167] S709. The algorithm layer returns the algorithm capability status.
[0168] Specifically, steps S708 and S709 can refer to the relevant descriptions of steps S408 - S409 in the corresponding embodiment above, and will not be elaborated here. Figure 4 Corresponding to the relevant descriptions of steps S408 - S409 in the corresponding embodiment, and will not be elaborated here.
[0169] S710, the CV pipeline determines the algorithm capability status.
[0170] The CV pipeline can determine whether it supports the semantic analysis algorithm capability based on the algorithm capability status. The specific method for determining the algorithm capability status can refer to the above Figure 4 The relevant description in step S409 of the corresponding embodiment is not repeated here.
[0171] If the CV pipeline determines that the electronic device does not support the algorithm capability of semantic analysis, the relevant process of the semantic analysis node of the video will not be executed.
[0172] If the CV pipeline determines that the electronic device supports the algorithmic capabilities of semantic analysis, it can further initialize the semantic model and load the video data. Figure 4 In the corresponding embodiments, the order of executing the initialization of the semantic model and the loading of the video data may not be distinguished. The CV pipeline may initialize the semantic model first, or load the video data first, or the process of initializing the semantic model and the process of loading the video data may be executed in parallel, which is not limited in the embodiments of the present application.
[0173] For ease of description, the following takes the process of loading a semantic model as an example.
[0174] S711. Load the semantic model.
[0175] S712. Return the semantic model loading status.
[0176] For specific steps S711 and S712, please refer to the above Figure 4 The relevant descriptions of step S411 to step S412 of the corresponding embodiment are not repeated here.
[0177] In a possible implementation, while loading the semantic model, the CV pipeline may also execute steps S713 to S715 to load video data from the media file library.
[0178] S713, the CV pipeline loads video data from the media file library.
[0179] S714. The CV pipeline loads a second number of videos from the media file library each time.
[0180] S715: The media file library returns the video data.
[0181] The CV pipeline can load a second quantity of videos from the media file library each time. In a possible implementation, the second quantity can be 10. This second quantity can be preset by the CV pipeline. For example, the second quantity can be an experimental test value. The embodiments of the present application do not limit the specific value of the second quantity. Among them, the value of the second quantity should be able to satisfy the timeliness of semantic search and not affect the smooth operation of the electronic device, etc.
[0182] It can be understood that when the CV pipeline loads a second quantity of videos from the media file library each time, it can first query whether there is video segment information corresponding to these videos in the video segment information library. If there is no video segment information corresponding to the video in the video segment information library, it means that the video has not been segmented, and then the video can be loaded, and the video segmentation processing flow from step S716 to step S720 can be executed. If there is video segment information corresponding to the video in the video segment information library, it means that the video has been segmented, and there is no need to segment the video again, and the process of generating video segment semantic vectors from step S721 to step S727 can be executed. In this way, the videos in the media file library will not be loaded repeatedly, and the process of video segmentation will not be executed repeatedly.
[0183] S716. The CV pipeline performs video slicing processing on the video.
[0184] The CV pipeline can slice a long video file to obtain multiple video slices.
[0185] In a possible implementation, the CV pipeline can regard a video with a duration greater than or equal to a first preset duration as a long video, and a video with a duration less than the first preset duration as a short video. Among them, the first preset duration can also be understood as a specification parameter of the long video. The first preset duration can be set by the CV pipeline. For example, the first preset duration can include 5 minutes. The embodiments of the present application do not limit the specific value of the first preset duration.
[0186] The CV pipeline can slice the long video at a second preset duration to obtain multiple video slices. Then the CV pipeline can transfer the multiple video slices to the algorithm layer to execute the video analysis algorithm. Among them, the second preset duration can be less than or equal to the first preset duration. The second preset duration can be set by the CV algorithm service. For example, the second preset duration can be 5 minutes or 3 minutes, etc. The embodiments of the present application do not limit the specific value of the second preset duration.
[0187] In a possible implementation, the cloud server can regularly push the value of the second preset duration to the electronic device. And the cloud server can configure different values of the second preset duration for different types of electronic devices.
[0188] Exemplarily, for an electronic device with a relatively fast processor processing speed and / or a relatively large memory space, the second preset duration pushed by the cloud server to the electronic device can be relatively long, such as 5 minutes, so that video slicing and segmentation and other processes can be quickly performed on the video. For an electronic device with a relatively slow processor processing speed and / or a relatively small memory space, the second preset duration pushed by the cloud server to the electronic device can be relatively short, such as 3 minutes, so as to reduce the lag phenomenon of the electronic device. The specific value of the second preset duration pushed by the cloud server to the electronic device is not limited in the embodiments of the present application. In this way, the cloud server can flexibly modify the value of the second preset duration for different electronic devices according to the actual situation, improving the flexibility and compatibility of the solution.
[0189] In another possible implementation, when the CV pipeline performs video slicing on a video, the resource situation of the codec can be considered. Exemplarily, if the resources of the codec are sufficient, the second preset duration can be relatively long. For example, the CV pipeline can slice the video with a duration of 5 minutes. If the resources of the codec are insufficient, the second preset duration can be relatively short. For example, the CV pipeline can slice the video with a duration of 3 minutes. This can reduce the resource occupancy of each video slice on the codec, so that the codec is not prone to lag when decoding the video.
[0190] Of course, the CV pipeline can also determine the value of the second preset duration according to the CPU occupancy rate, memory resource status, temperature, and / or power consumption of the electronic device, etc. The consideration of the value of the second preset duration is not limited in the embodiments of the present application.
[0191] S717. The CV pipeline calls the video analysis algorithm of the algorithm layer.
[0192] The CV pipeline can call the video analysis algorithm of the algorithm layer through the relevant interfaces of the algorithm adaptation layer, algorithm interface layer, and algorithm layer.
[0193] S718. The algorithm layer executes the video analysis algorithm.
[0194] The video analysis algorithm can include decoding video slices through a codec, extracting frames according to a frame extraction strategy, analyzing video frames according to an algorithm operation strategy, generating representative frames for video segmentation, etc.
[0195] In a possible implementation, the CV pipeline can transfer the path of the video slice, frame extraction strategy, algorithm operation strategy, etc. to the algorithm layer, and the algorithm layer can execute the video analysis algorithm on the video slice.
[0196] Among them, the frame extraction strategy can include extracting video frames according to key frames, can also include extracting video frames according to consecutive frames, and can also include other ways of extracting video frames, which are not limited in the embodiments of the present application.
[0197] Exemplarily, a key frame is also called an intra-coded picture frame (I-frame). A key frame can completely retain the image information of a frame of picture. If the frame extraction strategy is to extract video frames based on key frames, the algorithm layer can extract each key frame in the video as a video frame. As a video frame, a key frame can include more image information, can relatively completely represent the semantics of the video, and is convenient for the subsequent algorithm layer to analyze the video frame.
[0198] If the frame extraction strategy is to extract video frames based on consecutive frames, the algorithm layer can extract a video frame every preset step length, or can extract a video frame every preset time interval. It can be understood that the algorithm layer can extract video frames evenly or unevenly. For example, the algorithm layer can extract a video frame every 10 frames, or can extract a video frame every 1 second, or can extract a video frame every 10 frames in the first 5 seconds and then extract a video frame every 5 frames in the following, etc. The embodiments of the present application do not make any limitations.
[0199] In a possible implementation, when extracting video frames, factors such as the CPU resources, memory resources, temperature, and / or power consumption of the electronic device can be considered. Exemplarily, when the electronic device is in a situation of sufficient CPU resources, sufficient memory resources, low temperature, and / or low power consumption, etc., the density of video frame extraction by the algorithm layer can be relatively high, so that it is not easy to have frame loss, can extract more video frames, generate more semantic vectors, and then can more accurately implement video semantic search. When the electronic device is in a situation of insufficient CPU resources, insufficient memory resources, high temperature, and / or high power consumption, etc., the density of video frame extraction by the algorithm layer can be relatively low, so as to save the power consumption of the electronic device and reduce the lag of the electronic device, etc.
[0200] The algorithm operation strategy can be understood as the strategy adopted when performing key frame analysis on key frames. The algorithm operation strategy can include parallel analysis of key frames or serial analysis of key frames. Among them, key frame analysis can include analyzing video frames using a label detection algorithm, a clarity algorithm, a jitter algorithm, etc.
[0201] The clarity algorithm can be used to detect the clarity of each detailed shadow pattern and its boundary in a video frame. In a possible implementation, a clarity detection tool can be used to determine the clarity of multiple video frames in the video, and then compare the clarity of the video frame with the clarity of adjacent video frames to determine the clarity change value of the video frame.
[0202] The jitter algorithm can be used to detect the phenomenon of jitter or shake in the content displayed by video frames during video playback. It can also be understood that the jitter algorithm can be used to detect the smoothness of video frame data to evaluate image quality and video continuity. In possible implementations, video jitter detection methods such as the optical flow method based on image displacement, the feature point matching method, and the method based on the characteristics of image gray distribution can be used to detect the jitter degree of multiple video frames in the video.
[0203] If the algorithm running strategy is to perform parallel analysis on key frames, the algorithm layer can calculate the jitter, label detection, and / or clarity of the key frame in parallel. If the algorithm running strategy is to perform serial analysis on key frames, the algorithm layer can calculate the jitter, label detection, and / or clarity of the key frame serially, where the calculation order of jitter, label detection, and / or clarity is not limited.
[0204] In possible implementations, the algorithm running strategy can consider the CPU resources, memory resources, temperature, and / or power consumption of the electronic device, etc.
[0205] Exemplarily, when the electronic device is in a situation where the CPU resources are sufficient, the memory resources are sufficient, the temperature is relatively low, and / or the power consumption is relatively low, etc., the algorithm running strategy can be to perform parallel analysis on key frames, so that the analysis of key frames can be completed relatively quickly. When the electronic device is in a situation where the CPU resources are insufficient, the memory resources are insufficient, the temperature is relatively high, and / or the power consumption is relatively high, etc., the algorithm running strategy can be to perform serial analysis on key frames, so that the power consumption of the electronic device can be saved and the lag of the electronic device can be reduced, etc. Whether to perform parallel analysis or serial analysis on key frames specifically is not limited in the embodiments of the present application.
[0206] In this way, the image processing method of the embodiments of the present application can flexibly adjust the calculation method of key frame analysis according to the resources, temperature, and / or power consumption of the electronic device, etc., so that while considering the calculation speed, the power consumption of the electronic device and the situation of lag can be reduced as much as possible.
[0207] After analyzing each key frame, the algorithm layer can segment multiple key frames into video segments. It can be understood that when a video is played, the content it displays changes continuously as the video frames are played in sequence, but the degree of change is different. Therefore, the algorithm layer can divide key frames with similar content and small degree of change into the same video segment.
[0208] The algorithm layer can also select a representative frame from each video segment. Among them, the representative frame can be the starting frame, ending frame, middle frame, a random video frame, or the optimal frame with the highest score of the video segment, etc., which is not limited in the embodiments of the present application. It can be understood that video segmentation can facilitate the semantic model to analyze the video more accurately, thus facilitating user search.
[0209] S719. The algorithm layer returns video segment information.
[0210] The video segment information may include information such as the start time of the video segment, the end time of the video segment, the label information of the video segment, the jitter value of the video segment, the score of the video segment, and / or the quality data of the video segment, etc. Among them, the start time of the video segment and the end time of the video segment can be used to jump to the corresponding video frame when displaying the video search results for the user subsequently, improving the user experience. The label information of the video segment can be used to more accurately display the video search results when performing video search subsequently.
[0211] S720. The CV pipeline saves the video segment information.
[0212] The CV pipeline can save the video segment information into the video segment information library.
[0213] S721. Load the video segment information.
[0214] S722. Load the video segment information corresponding to the second quantity of videos.
[0215] Optionally, if the video segment information corresponding to the second quantity of videos has been generated in advance, the relevant processes of steps S713 - S720 may not be executed in the execution process of the corresponding embodiment. Figure 7 In the execution process of the corresponding embodiment, the relevant processes of steps S713 - S720 may not be executed.
[0216] S723. Return the video segment information.
[0217] S724. The CV pipeline judges the loading state of the semantic model.
[0218] The CV pipeline can judge whether the model is successfully loaded according to the model loading state. The specific judgment method of the model loading state can refer to the relevant description in step S712 above, and will not be elaborated here.
[0219] If the CV pipeline determines that the semantic model is not successfully loaded, the relevant processes of the semantic analysis node will not be continued.
[0220] If the CV pipeline determines that the semantic model is successfully loaded and the video segment information is successfully read, the semantic analysis algorithm can be executed on the video segment.
[0221] S725. The CV pipeline calls the semantic analysis algorithm.
[0222] After the CV pipeline reads the video segment information from the video segment information library, it can transfer the video segment information to the semantic model through the relevant interfaces such as the algorithm adaptation layer, the algorithm interface layer, and the algorithm layer. The semantic model can execute the semantic analysis algorithm to perform semantic analysis on the video segment.
[0223] S726. The semantic model returns the semantic vector of the video segment.
[0224] The semantic model can perform semantic analysis using representative frames. Exemplarily, the semantic model can input the representative frame into the image encoder of the CLIP model to obtain the semantic vector corresponding to the representative frame. It can be understood that a video includes multiple video segments, each video segment corresponds to a representative frame, and each representative frame corresponds to a semantic vector. Therefore, the semantic vector of the representative frame can also be understood as the semantic vector of the video segment, and a video can correspond to multiple video segment semantic vectors.
[0225] After the semantic model generates the semantic vector of the video segment, it can return the semantic vector of the video segment to the CV pipeline through relevant interfaces such as the algorithm adaptation layer, the algorithm interface layer, and the algorithm layer.
[0226] S727. Save the semantic vector of the video segment.
[0227] The CV pipeline can save the obtained semantic vector of the video segment into the semantic vector library.
[0228] It can be understood that in the semantic vector library, the semantic vector of the video segment can be saved in any possible way. In a possible implementation, the semantic vector of the video segment in the semantic vector library can be saved using a 768-bit array. The specific way of saving the semantic vector of the video segment in the semantic vector library is not limited in the embodiments of the present application.
[0229] Figure 8 The schematic diagram of the processing flow of video slicing and video segmentation such as the above steps S716 - S718 is shown. The process can include (1) long video slicing, (2) video segment segmentation, and (3) video segment semantic analysis.
[0230] (1) Long video slicing.
[0231] Before performing semantic analysis on the video, the long video file can be sliced first. As Figure 8 shown, the long video can be sliced into shorter videos with a duration of 5 minutes. The specific process of long video slicing can refer to the relevant description in the above step S716 and will not be elaborated here.
[0232] Slicing the long video can reduce the resource occupation of the codec and the system, making the video decoding by the codec smoother, and it is not easy for the electronic device to experience lag.
[0233] (2) Video segment segmentation.
[0234] The process of video segment segmentation can include video decoding, dynamic frame extraction, label detection, clarity detection, jitter detection, video segmentation, etc. As Figure 8As shown, the video slices can be extracted to obtain video frames, and the video frames can be analyzed to further segment the video and generate representative frames, for example, Figure 8 The three video segments in , each video segment corresponds to a representative frame. The specific process of video segmentation can refer to the relevant description in the above step S718, which will not be repeated here.
[0235] (3) Semantic analysis of video segments.
[0236] The semantic model can use the representative frame for semantic analysis. The process of semantic analysis of the specific video segment can refer to the relevant description in the above steps S725 and S726, which will not be repeated here. It can be understood that the semantic model can be deployed in the electronic device without the need for networking. When the electronic device is offline, a semantic vector can be generated for the picture or video, thereby improving the user experience.
[0237] (2) Construct an index of semantic vectors.
[0238] Figure 9 A timing diagram of index construction according to an embodiment of the present application is shown.
[0239] S901, charging and screen off trigger index building.
[0240] It is understandable that Figure 4 The analysis of the CV algorithm for charging and screen-off triggering in step S401 of the corresponding embodiment is similar. When the electronic device is charging and / or the screen is off, the gallery can trigger index construction, or the gallery can index construction based on the user's active trigger, or trigger index construction in other scenarios. The specific implementation method of triggering index construction is not limited in the embodiments of this application.
[0241] S902: The gallery starts index building service.
[0242] In the embodiment of the present application, the service for generating the semantic vector and the service for building the index can be different services. For example, the service for generating the semantic vector can be the first service, and the service for building the index can be the second service. The first service and the second service are different. The first service and the second service can be started and run at the same time, or they can be started and run one after another. The specific order of starting and running the first service and the second service is not limited in the embodiment of the present application.
[0243] S903: Load image information from the media file library and / or image information from the semantic vector library.
[0244] In a possible implementation, the file index proxy service may load image information in a media file library and image information in a semantic vector library.
[0245] After the index construction is triggered, the file index proxy service can load the image information in the media file library and pass the image information to the file index service engine of the index construction module. The file index service engine can build an index for the image based on the image tags and / or image names, etc., and save it in the media index library.
[0246] When the content in the semantic vector library is updated, the semantic vector library can send a message indicating that the semantic vector is updated to the file index proxy service. When the file index proxy service obtains the indication message, it can load the image information in the semantic vector library and pass the image information to the file index service engine. The file index service engine can build an index for the image based on the semantic vector and save it in the media index library.
[0247] It can be understood that the file index proxy service loading the image information in the media file library can be executed in the first process, and the file index proxy service loading the image information in the semantic vector library can be executed in the second process. The first process and the second process can be different processes. The first process and the second process can be started and run simultaneously, or can be started and run successively. The specific order of starting and running of the first process and the second process is not limited in the embodiments of the present application.
[0248] If the file index proxy service loads the image information from the media file library, it can quickly build an index for the image based on the image tags and / or image names, etc., so as to facilitate users to search and quickly provide the function of searching images for users; if the file index proxy service loads the image information from the semantic vector library, the built index can include the information of the image semantic vector. That is to say, building an index based on the media file library can quickly build an index for the image, and building an index based on the semantic vector library can make the image index contain more detailed image information. In this way, the image file desired by the user can be found faster and more accurately.
[0249] In another possible implementation, the file index proxy service can only load the image information in the media file library.
[0250] After the index construction is triggered, the file index proxy service can load the image information in the media file library and pass the image information to the file index service engine. The file index service engine can build a vector index for the image and save it in the media index library.
[0251] When the content in the semantic vector library is updated, the semantic vector library can send a message indicating to update the image file to the media file library, and the media file library updates the timestamp of the corresponding image based on the indication message.
[0252] When the file index proxy service obtains an image from the media file library, it can determine whether the image has been updated based on the timestamp of the image.
[0253] Exemplarily, taking a certain image as an example, the first timestamp of the image can be recorded in the media file library. After the file index service engine successfully constructs an index for the image, the second timestamp of the image can be recorded in the media index library, where the second timestamp is later than the first timestamp.
[0254] If, after the index of the image is constructed, the semantic algorithm module generates a semantic vector for the image, the semantic vector library can send a message to the media file library to indicate updating the image, and the media file library can update the first timestamp of the image to a third timestamp, where the third timestamp is later than the second timestamp.
[0255] When the file index proxy service obtains an image from the media file library and determines that the third timestamp of the image is later than the second timestamp, it indicates that after the index of the image is constructed, the media file library has updated the information of the image, and the index of the image needs to be reconstructed. Then the file index proxy service can pass the information of the image to the file index service engine. The file index proxy service can also obtain the semantic vector of the image from the semantic vector library according to the hash value of the image and pass the semantic vector to the file index service engine. The file index service engine can update the vector index of the image, save the updated image vector index to the media index library, and update the second timestamp of the image in the media index library to a fourth timestamp, where the fourth timestamp is later than the third timestamp.
[0256] It can be understood that the file index proxy service can also load image information in other ways. The specific implementation method of loading image information is not limited in the embodiments of the present application. For example, the file index proxy service can also only load the image information in the semantic vector library. When the content in the semantic vector library is updated, the semantic vector library can send a message to the file index proxy service to indicate that the semantic vector has been updated. When the file index proxy service obtains the indication message, it can load the image information in the semantic vector library and pass the image information to the file index service engine. The file index service engine can construct an index for the image based on the semantic vector and save it to the media index library.
[0257] S904. The file index proxy service calls the interface of the index construction module to construct an index.
[0258] S905. The index construction module constructs the vector index of the picture or video.
[0259] The index construction module can construct a vector index for an image or a video. The vector index can be understood as a record in a database. It can be understood that an image can correspond to one vector index, a video includes multiple video segments, and each video segment can correspond to one vector index. Then, a video can include multiple vector indexes, and different vector indexes correspond to different video segments in the video.
[0260] Among them, the vector index can include information such as the hash value of the image, the semantic vector of the image, the path of the image, the shooting time of the image, the shooting location of the image, and the timestamp of the image.
[0261] S906. Store the vector index in the media index library.
[0262] The index construction module can save the vector index to the media index library.
[0263] S907. The index construction ends.
[0264] After the index construction ends, the index construction module can return information indicating the end of the index construction to the image library. This information can include a first identifier indicating the end of the index construction, a second identifier indicating whether the index of a certain image is constructed successfully, and other information.
[0265] Among them, the first identifier can be of data types such as integer, boolean, string, etc. For example, when the first identifier is true, it can indicate that the index construction ends. The first identifier can also be other values. The specific value of the first identifier is not limited in the embodiments of the present application.
[0266] The second identifier can be of data types such as integer, boolean, string, etc. For example, when the second identifier of a certain image is true, it can indicate that the index of the image is constructed successfully. When the second identifier of a certain image is false, it can indicate that the index of the image is constructed failed. The second identifier can also be other values. The specific value of the second identifier is not limited in the embodiments of the present application.
[0267] It can be understood that if the index of a certain image is constructed successfully, the image library does not need to construct the index for this image anymore; if the index of a certain image is constructed failed, the image library can reconstruct the index for this image.
[0268] (3) Retrieve the text semantics.
[0269] Figure 10 Shows the timing diagram of the user's retrieval in the embodiments of the present application.
[0270] S1001. The user queries an image or a video.
[0271] Users can query pictures or videos by entering text in the search bar of the picture gallery. For example, users can enter search text in the search bar 302 as described above Figure 3 to query pictures or videos.
[0272] The search text is the text that describes the characteristics of the video required by the user. Exemplarily, the search text may include the shooting time of the image, the shooting location of the image, and the content of the picture shown in the image, etc. For example, the search text can be "Playing by the sea in City A last year". Regarding the specific content of the search text, this application does not make any limitations.
[0273] Optionally, users can also query pictures or videos by voice input. It can be understood that this application embodiment does not limit the way users instruct the picture gallery to query pictures or videos.
[0274] S1002. The picture gallery queries the text.
[0275] S1003. The file search proxy service calls the file search service engine to query pictures or videos.
[0276] The picture gallery can transfer the text information or voice information input by the user to the file search proxy service in the application function layer. Or, the picture gallery can transfer the keywords of the text information input by the user, or the keywords of the voice information to the file search proxy service.
[0277] The file search proxy service can call the file search service engine in the intelligent algorithm middle platform to query pictures or videos corresponding to the information input by the user.
[0278] S1004. Query pictures or videos.
[0279] S1005. The NLU extracts the text semantics.
[0280] Taking the search text "Playing by the sea in City A last year" as an example, the NLU semantic extraction module can divide the search statement into time, location, text semantics, etc. For example, the time can be "last year", the location can be "City A", and the text semantics can be "Playing by the sea".
[0281] S1006. The NLU semantic extraction module calls the text semantic analysis ability.
[0282] The NLU semantic extraction module can transfer the extracted text semantic part to the semantic algorithm module, so as to call the text semantic analysis ability of the semantic algorithm module to analyze the text semantics.
[0283] S1007. The semantic algorithm module performs text semantic analysis.
[0284] S1008. The semantic model generates a text semantic vector.
[0285] S1009. The semantic model returns the text semantic vector to the semantic algorithm module.
[0286] S1010. The semantic algorithm module returns the text semantic vector to the index construction module.
[0287] The semantic algorithm module can perform semantic analysis on the text semantics based on the semantic model, thereby generating a text semantic vector.
[0288] S1011. The index construction module queries the matching image index in the index according to the text semantic vector.
[0289] The index construction module can query the image semantic vector with a higher similarity to the text semantic vector in the media index library, where the image semantic vector can include a picture semantic vector and a video semantic vector. In a possible implementation, the vector similarity can be calculated through the cosine similarity calculation formula, or can be calculated through other methods, which is not limited in the embodiments of the present application.
[0290] If the similarity between the text semantic vector and the image semantic vector is greater than or equal to a certain preset threshold, it can be shown that the matching degree between the text semantic vector and the image semantic vector is relatively high and meets the search requirements. If the similarity between the text semantic vector and the image semantic vector is less than a certain preset threshold, it can be shown that the matching degree between the text semantic vector and the image semantic vector is relatively low and does not meet the search requirements. Among them, the preset threshold can be set in advance by the intelligent middle platform, and the specific value of the preset threshold is not limited in the embodiments of the present application.
[0291] S1012. The index construction module returns the image or video index to the picture library.
[0292] The index construction module can return the file index corresponding to the image file that meets the search requirements to the picture library. Among them, the file index can include information in the vector index of the image, such as the hash value of the image, the path of the image, etc.
[0293] S1013. The picture library queries the media file according to the picture index or video index.
[0294] The picture library can query the corresponding image file in the media file library according to information such as the hash value of the image and / or the path of the image.
[0295] S1014. The picture library loads the media file.
[0296] S1015. The electronic device presents the searched media file in the picture library interface.
[0297] The electronic device can display the pictures or videos found in the picture library to the user. Such asFigure 3 In the interface 301, the electronic device can display the searched pictures 303, pictures containing text 304, and videos 305. In a possible implementation, the cover of the video 305 can be displayed as a picture that satisfies the search result.
[0298] Optionally, the gallery can also generate a wonderful album for the user. For example, every once in a while, the gallery can select photos and videos of that period to generate a short review video. In some scenarios, a wonderful album can also be called a moment, a wonderful moment, a small moment, etc. The embodiment of the present application can also perform semantic analysis on the pictures or videos in the wonderful album one by one and generate a semantic vector. When a user queries for a picture or video, if a picture or video in the wonderful album meets the user's search requirements, the electronic device can display the wonderful album to the user. In a possible implementation, the cover of the wonderful album can display a picture that meets the search results.
[0299] In this way, the electronic device can match the image file that the user wants to find with the image files in the gallery based on the semantic model, thereby providing the user with images that are closer to the search content. The user can find the desired image file faster and more accurately.
[0300] Figure 11 A schematic diagram of a media semantic search framework provided in an embodiment of the present application is shown.
[0301] (1) CV algorithm service trigger.
[0302] With the above Figure 4 The analysis of the CV algorithm triggered by the gallery in step S401 of the corresponding embodiment is similar. When the electronic device is charging and / or the screen is off, the gallery can trigger the CV algorithm service, or the gallery can start the CV algorithm service according to the user's active trigger, or the CV algorithm service can be used in other scenarios. Figure 4 The description of step S401 of the picture library triggering the CV algorithm analysis in the corresponding embodiment will not be repeated here.
[0303] (2) Generate image semantic vector and construct image semantic index.
[0304] Exemplarily, the CV algorithm service can read image data in the media file library, wherein the image data can include information such as the bitmap of the image. The CV algorithm service can pass the image data to the semantic algorithm module of the intelligent algorithm module. The semantic algorithm module can use the semantic analysis algorithm to pre-process the image. For example, the pre-processing can include converting the image format into the image format required by the semantic model. The semantic algorithm module can also use the semantic model to generate a picture semantic vector for the image, and store the picture semantic vector in the semantic vector library. The specific process of generating the picture semantic vector can refer to the aboveFigure 4 The relevant descriptions of the corresponding embodiments will not be repeated.
[0305] The file index proxy service can load the picture information in the media file library and / or the picture information in the semantic vector library, and transfer the picture information to the file index service engine of the index construction module. The file index service engine can build an index for the picture and save the index to the media index library. The specific process of building the picture semantic index can refer to the above Figure 9 The relevant descriptions of the corresponding embodiments will not be repeated.
[0306] (3) Generate video semantic vectors and build video semantic indexes.
[0307] Exemplarily, the CV algorithm service can read the videos in the media file library and perform processes such as video slicing, video decoding, video frame extraction, video frame analysis, and video segmentation, so as to store the video segment information into the video segment information library. The semantic algorithm module can preprocess the video segments using semantic analysis algorithms, and can also use semantic models to generate the video segment semantic vectors of the video segments and store the video segment semantic vectors into the semantic vector library. The specific process of generating video semantic vectors can refer to the above Figure 7 The relevant descriptions of the corresponding embodiments will not be repeated.
[0308] The file index proxy service can load the video information in the media file library and / or the video information in the semantic vector library, and transfer the video information to the file index service engine of the index construction module. The file index service engine can build an index for the video and save the index to the media index library. The specific process of building the video semantic index can refer to the above Figure 9 The relevant descriptions of the corresponding embodiments will not be repeated.
[0309] (4) Semantic search.
[0310] When the user enters a search statement in the search box of the picture library, the picture library can call the NLU semantic extraction module integrated in the search engine to extract the text semantics, and the text semantics can generate text semantic vectors through the semantic model. The intelligent algorithm middle platform can search for matching image semantic vectors in the media index library according to the text semantic vectors and return the image files with a matching degree exceeding the preset threshold. In this way, the picture library can find the image files with a relatively high matching degree for the user. The specific process of semantic search can refer to the above Figure 10 The relevant descriptions of the corresponding embodiments will not be repeated.
[0311] The method of the embodiments of the present application will be described in detail below through specific embodiments. The following embodiments can be combined with each other or implemented independently, and the same or similar concepts or processes may not be repeated in some embodiments.
[0312] Figure 12 The figure shows an image processing method according to an embodiment of the present application. The method includes:
[0313] S1201. An electronic device displays a first interface of a first application, and the first interface includes an input box.
[0314] In an embodiment of the present application, the first application can be understood as an application that can view pictures or videos and has a search function. For example, the first application may include a gallery application.
[0315] The first interface can be understood as an interface in the first application that includes an input box. By entering text in the input box, the user can search for pictures or videos that match the entered text content. For example, the first interface may include the Figure 3 interface 301 in the corresponding embodiment above.
[0316] S1202. In response to an operation in which the user enters text in the input box, the electronic device displays a second interface, and the second interface includes one or more thumbnails.
[0317] In an embodiment of the present application, the second interface can be understood as an interface that can display search results, and thumbnails of pictures or videos that match the entered text content can be displayed in the interface. For example, the second interface may include the Figure 3 interface 301 in the corresponding embodiment above.
[0318] S1203. Among them, the thumbnails include thumbnails of videos, and the thumbnails of the videos are obtained by matching from an index library based on vectors of the input text. The index library stores video segment vectors corresponding to videos in the first application, and the video segment vectors corresponding to videos in the first application are obtained in advance by the following method.
[0319] In an embodiment of the present application, the index library can be understood as a data volume storing picture vectors or video segment vectors in the first application. For example, the index library may include the Figure 9 media index library in the corresponding embodiment above.
[0320] The video segment vector can be understood as the Figure 7 semantic vector of the video segment in the corresponding embodiment above.
[0321] S1204. The videos in the first application are segmented for the first time according to a preset duration to obtain video segments after the first segmentation.
[0322] In an embodiment of the present application, the preset duration can be understood as the duration for segmenting the video for the first time. For example, the preset duration may be the Figure 7 second preset duration in the corresponding embodiment above, which will not be elaborated here.
[0323] The process of specifically obtaining the video segments after the first segmentation can refer to the above Figure 7 description of the CV pipeline in step S716 of the corresponding embodiment for video slicing of the video, which will not be elaborated here.
[0324] S1205. Perform a second video segmentation on the video segments after the first segmentation according to the image similarity of adjacent frames to obtain the video segments after the second segmentation.
[0325] In the embodiment of the present application, the process of specifically obtaining the video segments after the second segmentation can refer to the above Figure 7 description of the algorithm layer in step S718 of the corresponding embodiment for executing the video analysis algorithm, which will not be elaborated here.
[0326] S1206. Generate video segment vectors for the video segments after the second segmentation respectively.
[0327] In the embodiment of the present application, the process of specifically generating the video segment vectors can refer to the above Figure 7 description in steps S725 and S726 of the corresponding embodiment, which will not be elaborated here.
[0328] The electronic device can match the image file that the user wants to search for with the image files in the image library based on the text vector and the video segment vectors, so as to provide the user with an image closer to the search content. In this way, the image file that the user wants can be found faster and more accurately through vector similarity analysis.
[0329] Optionally, before Figure 12 performing a second video segmentation on the video segments after the first segmentation according to the image similarity of adjacent frames based on the corresponding embodiment, it may further include: decoding the video segments after the first segmentation to obtain the decoded video segments; extracting one or more video frames from the decoded video segments; performing a second video segmentation on the video segments after the first segmentation according to the image similarity of adjacent frames may include: performing video frame analysis on one or more video frames to obtain the image similarity of adjacent video frames; if the image similarity of adjacent video frames meets the similarity threshold, dividing the adjacent video frames into the same video segment, and if the image similarity of adjacent video frames does not meet the similarity threshold, dividing the adjacent video frames into different video segments.
[0330] In the embodiment of the present application, the processes of decoding the video segments after the first segmentation, extracting one or more video frames from the decoded video segments, performing video frame analysis on one or more video frames, and performing the second video segmentation, etc. can refer to the above Figure 7 description in step S718 of the corresponding embodiment, which will not be elaborated here.
[0331] It can be understood that the image similarity meeting the similarity threshold means that when the image similarity is greater than or equal to the similarity threshold, the similarity of adjacent video frames is relatively high, and when the image similarity is less than the similarity threshold, the similarity of adjacent video frames is relatively low; it can also be understood that when the image similarity is less than or equal to the similarity threshold, the similarity of adjacent video frames is relatively high, and when the image similarity is greater than the similarity threshold, the similarity of adjacent video frames is relatively low. The embodiments of the present application do not make a limitation in this regard.
[0332] In this way, the video can be segmented, which is convenient for the semantic model to analyze the video more accurately and convenient for users to search.
[0333] Optionally, on the basis of the Figure 12 corresponding embodiment, extracting one or more video frames from the decoded video segment may include: extracting one or more video frames from the decoded video segment according to a frame extraction strategy, and the frame extraction strategy includes: extracting one or more video frames from the decoded video segment according to key frames, or extracting one or more video frames from the decoded video segment according to consecutive frames.
[0334] In the embodiments of the present application, the frame extraction strategy may refer to the relevant description of the frame extraction strategy in step S718 in the above Figure 7 corresponding embodiment, and will not be elaborated here.
[0335] The video frame may include image information, which can represent the semantics of the video, facilitating subsequent analysis of the video frame and thus more accurately implementing video semantic search.
[0336] Optionally, on the basis of the Figure 12 corresponding embodiment, video frame analysis may include one or more of the following: label detection, clarity detection, or jitter detection.
[0337] In the embodiments of the present application, the video frame analysis may refer to the relevant description in the algorithm operation strategy in step S718 in the above Figure 7 corresponding embodiment, and will not be elaborated here.
[0338] Performing video frame analysis based on the algorithm operation strategy can flexibly adjust the calculation method of video frame analysis, so that while considering the calculation speed, the power consumption of the electronic device and the occurrence of jamming can be reduced as much as possible.
[0339] Optionally, on the basis of the Figure 12 corresponding embodiment, after responding to the operation of the user inputting text in the input box, it may further include: generating a text vector for the text; the electronic device displays a second interface, and the second interface includes one or more thumbnails, which may include: determining the similarity between the text vector and the video segment vector; if the similarity between the text vector and the video segment vector is greater than or equal to a first threshold, the electronic device displays a second interface, and the second interface includes thumbnails of the video.
[0340] In the embodiments of the present application, the text vector can be understood as the text semantic vector in the corresponding above-mentioned Figure 10 embodiment. The first threshold can be understood as the preset threshold in the corresponding above-mentioned Figure 10 embodiment.
[0341] For the process of determining the similarity between the text vector and the video segment vector, reference can be made to the relevant description of step S1011 in the corresponding above-mentioned Figure 10 embodiment, which will not be elaborated herein.
[0342] The electronic device can match the image file that the user wants to search for with the image files in the image library based on the semantic vector, so as to provide the user with an image closer to the search content. The user can find the desired image file faster and more accurately.
[0343] Optionally, on the basis of the corresponding Figure 12 embodiment, before generating video segment vectors for the video segments after the second segmentation, it may further include: saving the relevant information of the video to the index library, where the relevant information of the video includes one or more of the following: the hash value of the video, the path of the video, or the first timestamp of the video; after generating video segment vectors for the video segments after the second segmentation, it may further include: sending a first instruction to the file library storing the videos in the first application; the file library updates the timestamp of the video to the second timestamp based on the first instruction; determining that the second timestamp is later than the first timestamp, updating the relevant information of the video in the index library, where the relevant information of the video includes one or more of the following: the hash value of the video, the video segment vector, the path of the video, or the third timestamp of the video, and the third timestamp is later than the second timestamp.
[0344] In the embodiments of the present application, the file library storing the videos in the first application can be understood as the media file library in the corresponding above-mentioned Figure 9 embodiment.
[0345] The first instruction can be used to indicate a message for updating the image file, and the media file library updates the timestamp of the corresponding image based on the indication message.
[0346] The first timestamp can be understood as the second timestamp in the media index library in the corresponding above-mentioned Figure 9 embodiment; the second timestamp can be understood as the third timestamp in the media file library in the corresponding above-mentioned Figure 9 embodiment; the third timestamp can be understood as the fourth timestamp in the media index library in the corresponding above-mentioned Figure 9 embodiment.
[0347] For the process of updating the relevant information of the video in the index library, reference can be made to the above-mentioned Figure 9For the relevant description in another possible implementation in step S903 of the corresponding embodiment, it will not be elaborated here.
[0348] By updating the timestamp information in the media index library and the media file library in a timely manner, the analysis status of the current video can be accurately judged, so that the process of generating semantic vectors and indexes for a certain video will not be repeated, and thus the image file desired by the user can be found faster and more accurately.
[0349] Optionally, on the basis of the Figure 12 corresponding embodiment, before generating video segment vectors for the video segments after the second segmentation, it may further include: saving the relevant information of the video to the index library, where the relevant information of the video includes one or more of the following: the hash value of the video, the path of the video, or the first timestamp of the video; after generating video segment vectors for the video segments after the second segmentation, it may further include: sending a first instruction to the index library; based on the first instruction, the index library updates the relevant information of the video, where the relevant information of the video includes one or more of the following: the hash value of the video, the video segment vector, the path of the video, or the second timestamp of the video, and the second timestamp is later than the first timestamp.
[0350] In the embodiments of the present application, saving the relevant information of the video to the index library can be understood as loading the image information from the media file library, and the index library updating the relevant information of the video based on the first instruction can be understood as loading the image information from the semantic vector library. The specific process of updating the relevant information of the video in the index library can refer to the Figure 9 relevant description in a possible implementation in step S903 of the corresponding embodiment, and will not be elaborated here.
[0351] The image information loaded from the media file library can quickly build an index for the image based on the image tags and / or image names, etc., so as to quickly provide the function of searching for images for users; the image information loaded from the semantic vector library, then the index constructed can include the information of the image semantic vector, so that the image index contains more detailed image information, and in this way, the image file desired by the user can be found faster and more accurately.
[0352] Optionally, on the basis of the Figure 12 corresponding embodiment, the preset duration is obtained from the cloud server or determined based on the resources of the codec.
[0353] In the embodiments of the present application, the method for determining the preset duration can refer to the Figure 7 relevant description in step S716 of the corresponding embodiment, and will not be elaborated here.
[0354] It can be understood that by obtaining the preset duration from the cloud server, the cloud server can flexibly modify the value of the second preset duration for different electronic devices according to the actual situation, which improves the flexibility and compatibility of the solution. Determining the preset duration based on the resources of the codec can reduce the resource occupancy of each video slice on the codec, so that the codec is not prone to jamming when decoding the video.
[0355] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0356] The above mainly introduces the solution provided by the embodiments of this application from the perspective of the method. To implement the above functions, it includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, combined with the method steps of each example described in the embodiments disclosed in this article, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0357] The embodiments of this application can divide the functional modules of the device for implementing the method according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of this application is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0358] As Figure 13 shown is a schematic structural diagram of a chip provided by an embodiment of this application. The chip 1300 includes one or more than two (including two) processors 1301, a communication line 1302, a communication interface 1303, and a memory 1304.
[0359] In some embodiments, the memory 1304 stores the following elements: executable modules or data structures, or subsets thereof, or extended sets thereof.
[0360] The methods described in the embodiments of the present application above can be applied to or implemented by the processor 1301. The processor 1301 may be an integrated circuit chip with the ability to process signals. During implementation, the steps of the above methods can be completed by the integrated logic circuit in hardware or instructions in software form in the processor 1301. The above-mentioned processor 1301 may be a general-purpose processor (e.g., a microprocessor or a conventional processor), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate, transistor logic devices, or discrete hardware components. The processor 1301 can implement or execute various processing-related methods, steps, and logic block diagrams disclosed in the embodiments of the present application.
[0361] The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or implemented and completed by a combination of hardware and software modules in the decoding processor. Among them, the software module can be located in a mature storage medium in the art such as a random access memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable read-only memory (EEPROM). This storage medium is located in the memory 1304, and the processor 1301 reads the information in the memory 1304 and combines its hardware to complete the steps of the above method.
[0362] The processor 1301, the memory 1304, and the communication interface 1303 can communicate with each other through the communication line 1302.
[0363] In the above embodiments, the instructions stored in the memory for the processor to execute can be implemented in the form of a computer program product. Among them, the computer program product can be pre-written in the memory in advance, or downloaded and installed in the memory in software form.
[0364] Embodiments of the present application also provide a computer program product including one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server or a data center including one or more available media integrated. For example, the available medium may include magnetic media (such as floppy disks, hard disks, or magnetic tapes), optical media (such as digital versatile discs (DVDs)), or semiconductor media (such as solid state disks (SSDs)).
[0365] Embodiments of the present application also provide a computer-readable storage medium. The methods described in the above embodiments may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. The computer-readable medium may include computer storage media and communication media, and may also include any medium that can transfer a computer program from one place to another. The storage medium may be any target medium accessible by a computer.
[0366] As a possible design, the computer-readable medium may include a compact disc read-only memory (CD-ROM), RAM, ROM, EEPROM, or other optical disc memories; the computer-readable medium may include magnetic disk memories or other magnetic disk storage devices. Moreover, any connection line may also be appropriately referred to as a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies (such as infrared, radio, and microwave), then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, magnetic disks and optical discs include optical discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where magnetic disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers.
[0367] Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processing unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processing unit of the computer or other programmable data processing device generate means for implementing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
Claims
1. An image processing method, characterized in that, the method includes: An electronic device displays a first interface of a first application, and the first interface includes an input box; In response to a user's operation of inputting text in the input box, the electronic device displays a second interface, and the second interface includes one or more thumbnails; Wherein, the thumbnails include thumbnails of videos, and the thumbnails of the videos are obtained by matching from an index library based on the vector of the input text. The index library stores video segment vectors corresponding to the videos in the first application, and the video segment vectors corresponding to the videos in the first application are obtained in advance by the following method: Perform a first video segmentation on the videos in the first application according to a preset duration to obtain video segments after the first segmentation; Perform a second video segmentation on the video segments after the first segmentation according to the image similarity of adjacent frames to obtain video segments after the second segmentation; Generate video segment vectors for the video segments after the second segmentation respectively.
2. The method according to claim 1, characterized in that, Before performing the second video segmentation on the video segments after the first segmentation according to the image similarity of adjacent frames, it further includes: Decode the video segments after the first segmentation to obtain decoded video segments; Extract one or more video frames from the decoded video segments; Performing a second video segmentation on the video segments after the first segmentation according to the image similarity of adjacent frames includes: Perform video frame analysis on the one or more video frames to obtain the image similarity of adjacent video frames; If the image similarity of adjacent video frames meets the similarity threshold, the adjacent video frames are divided into the same video segment. If the image similarity of adjacent video frames does not meet the similarity threshold, the adjacent video frames are divided into different video segments.
3. The method according to claim 2, characterized in that, Extracting one or more video frames from the decoded video segments includes: Extract one or more video frames from the decoded video segments according to a frame extraction strategy, and the frame extraction strategy includes: Extract one or more video frames from the decoded video segments according to key frames, or extract one or more video frames from the decoded video segments according to consecutive frames.
4. The method according to claim 2 or 3, characterized in that, The video frame analysis includes one or more of the following: label detection, clarity detection, or jitter detection.
5. The method according to any one of claims 1-4, characterized in that, After the operation of responding to the user's input of text in the input box, it further includes: Generate a text vector for the text; The electronic device displays a second interface, and the second interface includes one or more thumbnails, including: Judge the similarity between the text vector and the video segment vector; If the similarity between the text vector and the video segment vector is greater than or equal to a first threshold, the electronic device displays the second interface, and the second interface includes the thumbnail of the video.
6. The method according to any one of claims 1-5, characterized in that, Before generating video segment vectors for the video segments after the second segmentation, the following steps are further included: Save the relevant information of the video to the index library, where the relevant information of the video includes one or more of the following: the hash value of the video, the path of the video, or the first timestamp of the video; After generating video segment vectors for the video segments after the second segmentation, the following steps are further included: Send a first instruction to the file library storing the video in the first application; Based on the first instruction, the file library updates the timestamp of the video to a second timestamp; Determine that the second timestamp is later than the first timestamp, and update the relevant information of the video in the index library, where the relevant information of the video includes one or more of the following: the hash value of the video, the video segment vector, the path of the video, or the third timestamp of the video, and the third timestamp is later than the second timestamp.
7. The method according to any one of claims 1-5, wherein, Before generating video segment vectors for the video segments after the second segmentation, the following steps are further included: Save the relevant information of the video to the index library, where the relevant information of the video includes one or more of the following: the hash value of the video, the path of the video, or the first timestamp of the video; After generating video segment vectors for the video segments after the second segmentation, the following steps are further included: Send a first instruction to the index library; Based on the first instruction, the index library updates the relevant information of the video, where the relevant information of the video includes one or more of the following: the hash value of the video, the video segment vector, the path of the video, or the second timestamp of the video, and the second timestamp is later than the first timestamp.
8. The method according to any one of claims 1-7, wherein, The preset duration is obtained from the cloud-side server or determined based on the resources of the codec.
9. An electronic device, wherein, It includes: A memory and a processor, the memory is used to store a computer program, and the processor is used to execute the computer program to perform the method according to any one of claims 1-8.
10. A computer-readable storage medium, wherein, The computer-readable storage medium stores instructions, and when the instructions are executed, the computer is made to execute the method according to any one of claims 1-8.
11. A computer program product, wherein, It includes a computer program, and when the computer program is run, the electronic device is made to execute the method according to any one of claims 1-8.
Citation Information
Patent Citations
Visual media personalized search method and device
CN113641857A
Text and video mutual inspection method and device, equipment and storage medium
CN115438169A
Timestamp-based ElasticSearch index creating and updating system and method
CN116383216A
Method and device for constructing search index database and search method and device
CN117033695A
Cited By
Information collection method and device, electronic equipment, storage medium and program product
CN121681971A