Method, apparatus and electronic device for previewing multi-modal content in AI data lake
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]在AI数据湖中,通过多模态数据集承载多模态(例如音频模态或者视频模态)内容,而由于多模态数据集本身是一种容器,而非单个可以直接被播放或者预览的文件,因此,对于客户端或者浏览器而言,无法将多模态数据集直接理解为可播放或者预览的文件进行渲染,从而,如何播放或者预览多模态数据集中的多模态内容是当前面临的技术问题
[0008]通过上述方法,通过获得第一字段的虚拟统一资源定位符,并基于虚拟统一资源定位符的属性信息,构造获得请求,进一步地,通过解析获得请求中的虚拟统一资源定位符以获得用于确定第一物理地址信息的第一元数据,并基于第一物理地址信息获得第一多模态内容,从而基于第一多媒体内容构造的响应实现第一多模态内容的预览。在实现预览的过程中,将第一字段的预览操作所对应的预览请求转换为基于第一物理地址信息获得相应第一多模态内容的二进制数据的获得请求,并基于获得的第一多模态内容的二进制数据构造相应的响应,在收到响应后,原生渲染组件能够识别该响应,无需得知响应中的数据来自多模态数据集,且在该过程中,无需改变多模态数据集的第一预设格式、无需将第一字段导出为独立的多模态内容或者无需将第一字段反序列化等,从而通过低开销的方式实现多模态数据集中的指定的字段所映射的多模态内容的预览。
Smart Images

Figure CN122549409A_ABST
Abstract
Description
Technical Field
[0001] This article relates to the field of computer technology, and more specifically, to a method, apparatus, and electronic device for previewing multimodal content in an AI (Artificial Intelligence) data lake. Background Technology
[0002] An AI (Artificial Intelligence) data lake can be a unified data infrastructure that supports AI / ML (Machine Learning) optimization.
[0003] In AI data lakes, multimodal datasets are used to carry multimodal content (such as audio or video modalities). However, since a multimodal dataset is a container rather than a single file that can be directly played or previewed, clients or browsers cannot directly interpret a multimodal dataset as a playable or previewable file for rendering. Therefore, how to play or preview the multimodal content in a multimodal dataset is a current technical problem. Summary of the Invention
[0004] This content section is provided to briefly introduce the ideas, which will be described in detail in the examples section later. This content section is not intended to identify key or essential features of the claimed content, nor is it intended to limit the scope of the claimed content.
[0005] Firstly, a method for previewing multimodal content in an AI data lake is provided, including: In response to the preview operation of the first field, a virtual uniform resource locator (VURL) for the first field is obtained. The first field belongs to a field in the multimodal dataset. The first field is used to map the first multimodal content. The multimodal dataset is located in the AI data lake. The AI data lake presents the multimodal dataset in a first preset format. The first multimodal content belongs to any second multimodal content in the multimodal dataset. The VURL carries the attribute information of the first field. Based on the attribute information of the virtual uniform resource locator, a request is constructed; Parse the Virtual Uniform Resource Locator in the request and obtain the first metadata based on the parsed attribute information; Based on the first metadata, obtain the first physical address information; Based on the first physical address information, the first multimodal content is obtained, and a response to the request is constructed based on the obtained first multimodal content. The response is used to preview the first multimodal content.
[0006] Secondly, a preview device for multimodal content in an AI data lake is provided, comprising: The first acquisition module is used to obtain the virtual unified resource locator of the first field in response to the preview operation of the first field. The first field belongs to a field in the multimodal dataset. The first field is used to map the first multimodal content. The multimodal dataset is located in the AI data lake. The AI data lake presents the multimodal dataset in a first preset format. The first multimodal content belongs to any second multimodal content in the multimodal dataset. The virtual unified resource locator carries the attribute information of the first field. The first construction module is used to construct an acquisition request based on the attribute information of the virtual uniform resource locator; The parsing module is used to parse the Virtual Uniform Resource Locator in the request and obtain the first metadata based on the parsed attribute information; The second obtaining module is used to obtain the first physical address information based on the first metadata; The second construction module is used to obtain the first multimodal content based on the first physical address information, and to construct a response to the acquisition request based on the obtained first multimodal content, wherein the response is used to preview the first multimodal content.
[0007] Thirdly, an electronic device is provided, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method in the first method aspect.
[0008] The above method obtains the Virtual Uniform Resource Locator (VURI) of the first field and constructs a request based on the attribute information of the VURI. Further, it parses the VURI in the request to obtain first metadata used to determine the first physical address information, and obtains the first multimodal content based on the first physical address information. This allows for a preview of the first multimodal content based on the response constructed from the first multimedia content. During the preview process, the preview request corresponding to the preview operation of the first field is converted into a request to obtain binary data of the corresponding first multimodal content based on the first physical address information. A corresponding response is constructed based on the obtained binary data of the first multimodal content. Upon receiving the response, the native rendering component can recognize it without needing to know that the data in the response comes from the multimodal dataset. Furthermore, this process does not require changing the first preset format of the multimodal dataset, exporting the first field as independent multimodal content, or deserializing the first field. Therefore, it achieves a low-overhead preview of the multimodal content mapped to a specified field in the multimodal dataset.
[0009] Other features and advantages will be described in detail in the following examples section. Attached Figure Description
[0010] The above and other features, advantages, and aspects of this document will become more apparent when viewed in conjunction with the accompanying drawings and the following examples. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 This is a schematic diagram illustrating the implementation environment of a preview method for multimodal content in an AI data lake, based on certain scenarios. Figure 2 This is a flowchart illustrating a preview method for multimodal content in an AI data lake, based on various scenarios. Figure 3 This is another schematic diagram illustrating the implementation environment of a preview method for multimodal content in an AI data lake, based on certain scenarios. Figure 4 This is a block diagram of a preview device for multimodal content in an AI data lake, shown according to certain scenarios; Figure 5 These are schematic diagrams of the structure of electronic devices shown in various scenarios. Detailed Implementation
[0011] The following description will be given in more detail with reference to the accompanying drawings. While certain scenarios are shown in the drawings, it should be understood that this document can be implemented in various forms and should not be construed as limited to the scenarios described herein. Rather, these scenarios are provided to provide a more thorough and complete understanding of this document. It should be understood that the accompanying drawings and the scenarios depicted are for illustrative purposes only and are not intended to limit the scope of this document.
[0012] It should be understood that the steps described in the method may be performed in different orders and / or in parallel. Furthermore, the method may include additional steps and / or omit the steps shown. The scope of this document is not limited in this respect.
[0013] The term "comprising" and its variations can be open-ended, meaning "including but not limited to". The term "based on" can mean "at least partially based on". The term "one case" means "at least one case"; the term "another case" means "at least one additional case"; the term "some cases" means "at least some cases". Definitions of other terms will be given in the following description.
[0014] It should be noted that the concepts of "first" and "second" are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.
[0015] It should be noted that the modifiers “one” and “multiple” can be illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as “one or more”.
[0016] The names of messages or information exchanged between multiple devices are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0017] It is understandable that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant regulations.
[0018] As the background technology indicates, in AI data lakes, multimodal datasets are used to carry multimodal content. For example, multimedia content is typically stored as a field in the multimodal dataset in the form of Blob (Binary Large Object), Binary, or LargeBinary, and together with metadata describing the multimodal dataset, they constitute a multimodal dataset. However, since a multimodal dataset is a container rather than a single file that can be directly played or previewed, clients or browsers cannot directly interpret it as a playable or previewable file for rendering. Therefore, how to play or preview the multimodal content within a multimodal dataset is one of the key research focuses currently.
[0019] Browsers or clients cannot read, understand, or parse the underlying structure of multimodal datasets. Therefore, in related technologies, the following methods are typically used to preview the multimodal content mapped to fields in a multimodal dataset: First, the fields used to map multimodal content in the multimodal dataset are exported in advance as independent images, audio, video or files, and then the address of the exported file is provided to the module that implements playback or preview (such as a browser). Second, the fields in the multimodal dataset used to map multimodal content are deserialized to obtain the multimedia content, and then the multimedia content is returned to the module that implements playback or preview. Third, convert the multimodal dataset into a file suitable for playback or preview.
[0020] These preview solutions introduce additional copies of multimodal content, processing links, and synchronization issues, resulting in significant overhead for previewing. Furthermore, they can easily lead to high bandwidth consumption and long initial screen latency in scenarios involving large files (i.e., multimodal content) and batch previews.
[0021] In view of this, this paper provides a method, device, storage medium, electronic device and program product for previewing multimodal content in an AI data lake. The following explanation and description are provided in conjunction with the accompanying drawings.
[0022] Figure 1 These are schematic diagrams illustrating application environments under certain circumstances, for reference only. Figure 1 Client 101 can provide an interactive interface for user 102, which allows user 102 to trigger preview operations and preview multimodal content. Client 101 can be, for example, a mobile phone or a laptop.
[0023] Figure 2This is a flowchart illustrating a method for previewing multimodal content in an AI data lake, based on various scenarios. This method can be implemented by a client, a browser, a system consisting of a client, a browser, and a server, or by other software or hardware capable of previewing multimodal content in an AI data lake. This document does not impose any limitations on this method.
[0024] Reference Figure 2 The preview method for multimodal content in the AI data lake mentioned above may include steps S210, S220, S230, S240 and S250.
[0025] In step S210, in response to the preview operation of the first field, a virtual Uniform Resource Locator (VURL) for the first field is obtained. The first field belongs to a field in the multimodal dataset. The first field is used to map the first multimodal content. The multimodal dataset is located in the AI data lake. The AI data lake presents the multimodal dataset in a first preset format. The first multimodal content belongs to any second multimodal content in the multimodal dataset. The VURL carries the attribute information of the first field.
[0026] The first field can be a field storing multimodal content in Blob format, and it is a field within the multimodal dataset. The multimodal dataset can be a dataset presented in an AI data lake using a first preset format. For example, the first preset format can refer to a format that includes metadata for locating physical address information, such as columnar or containerized formats. A multimodal dataset could be a columnar Lance file. As another example, the first preset format can refer to a format that includes metadata for locating physical address information and supports fragmented reading. Explanations and instructions regarding fragmented reading can be found below.
[0027] Taking a columnar format as an example, a multimodal dataset can contain multiple columns. Each row in the multimodal dataset can represent a single sample data point. These sample data points can support tasks such as model training and inference. A model can be a large multimodal model, which refers to an artificial intelligence model capable of simultaneously understanding, processing, and generating multiple modalities. One column in the multimodal dataset can correspond to the first field, and the remaining columns can correspond to the second, third, etc., fields. The second field can be, for example, the sample identifier (or row number), and the third field can be, for example, the sample's classification label. Of course, a multimodal dataset can also include fields with other meanings; this paper does not limit this.
[0028] As an example, the client can provide an interactive interface that displays a table corresponding to the multimodal dataset. The table may include the field names corresponding to the first field, the second field, and the third field mentioned above. Users can trigger a preview operation for the first field by interacting with the field name of the first field in a row of the table (e.g., clicking). The preview operation can correspond to a preview request, and the preview request carries the virtual Uniform Resource Locator (URL) of the first field.
[0029] The first field is not data that directly represents the first multimodal content and cannot be previewed directly. It can be understood as a carrier storing the first multimodal content using raw 0 / 1 binary data, containing attribute information such as the identifier of the multimodal dataset, the version snapshot of the multimodal dataset, the sample identifier, the name of the first field, and the format type of the multimodal content (e.g., image / png, video / mp4, application / pdf, etc.). Based on the interacted first field, its attribute information can be obtained, thus constructing a virtual Uniform Resource Locator (URL). Furthermore, the format type of the multimodal content can be explicitly declared in the URL or inferred from other information, such as the object's metadata. For an explanation and description of the object, please refer to the following content.
[0030] A Uniform Resource Locator (URL), commonly known as a web address, is a standard string used to uniquely identify and locate any network resource. It tells the browser / client where the resource is located, what protocol to use to access it, and which file to retrieve. However, the virtual URL discussed in this article is merely an identifier; the backend / object storage does not contain a corresponding actual file. The virtual URL can shield the underlying structure of multimodal datasets, allowing browser-native components such as images, videos, and audio to directly preview and render multimedia without needing to adapt to the underlying structure of the multimodal dataset.
[0031] In step S220, a request is constructed based on the attribute information of the virtual Uniform Resource Locator.
[0032] As can be seen from the above, the attribute information may include the identifier of the multimodal dataset, the version snapshot of the multimodal dataset, the sample identifier, the name of the first field, and the format type of the multimodal content, etc. The request is used to request the binary data of the first multimodal content based on these attribute information.
[0033] In step S230, the virtual uniform resource locator in the request is parsed, and the first metadata is obtained based on the parsed attribute information.
[0034] The first metadata is used to characterize the physical address that can locate the first multimodal content.
[0035] In step S240, the first physical address information is obtained based on the first metadata.
[0036] The first physical address information may include the following: The object containing the first multimodal content includes the multimodal dataset or an associated object of the multimodal dataset; The starting offset of the first multimodal content; The byte length of the first multimodal content; The format type of the first multimodal content.
[0037] If the first field is stored in an associated object within the multimodal dataset (e.g., external object storage), the address of the associated object needs to be determined. If the first field is embedded within the multimodal dataset, the address within the multimodal dataset containing the first field needs to be determined (e.g., a fragmented file as described below). Depending on the size of the first field, you can choose to store it in an associated object within the multimodal dataset or embed it within the multimodal dataset. For example, when the first field is small, it can be embedded within the multimodal dataset; conversely, when the first field is large, it can be stored in an associated object within the multimodal dataset.
[0038] In step S250, the first multimodal content is obtained based on the first physical address information, and a response to the request is constructed based on the obtained first multimodal content. The response is used to preview the first multimodal content.
[0039] Understandably, the starting position for reading is determined based on the initial offset of the first multimodal content, and the ending position for reading is determined by combining this with the byte length of the first multimodal content. The starting and ending positions determine the range of bytes read. The format type of the first multimodal content (e.g., PNG) can be used to instruct the client / browser how to render (i.e., which native rendering component to select).
[0040] It should be noted that the binary data of the first multimodal content can be obtained based on the first physical address information. This binary data is then encapsulated to obtain a media resource (i.e., the first multimodal content) that can be recognized by the browser / client. This media resource can be carried in the aforementioned response. The response here can refer to a standard or streaming response to an HTTP (Hypertext Transfer Protocol) request from the client / browser. The native rendering component in the client / browser can recognize this response and, based on it, preview the first multimodal content.
[0041] Thus, by using the above method, after receiving the response, the browser / client can recognize the response without knowing that the data in the response comes from the multimodal dataset. That is, the browser / client does not need to understand the underlying results of the multimodal dataset. Moreover, in this process, there is no need to change the first preset format of the multimodal dataset, export the first field as independent multimodal content, or deserialize the first field, etc., thereby enabling a preview of the multimodal content mapped by the specified field in the multimodal dataset in a low-overhead manner.
[0042] In some cases, the step of obtaining the virtual Uniform Resource Locator (VURL) of the first field in response to the preview operation of the first field may include: constructing the VURL of the first field according to a second preset format in response to the preview operation of the first field, so as to obtain the VURL of the first field. The second preset format includes multiple formats, and different second preset formats use different data organization methods to encapsulate the attribute information of the first field into the VURL.
[0043] Among the various second preset formats, one can be randomly selected as the data organization method on which the virtual uniform resource locator is constructed, or one can be selected according to the actual preview scenario. This article does not limit this choice.
[0044] The above data organization methods can include the following types: The first type is used to represent the attribute information of the first field carried in the path parameter; The second type is used to represent the attribute information of the first field carried in the query parameters; The third type is used to represent the attribute information of the first field carried in the encryption token; The fourth type is used to represent the attribute information of the first field carried in the signature token.
[0045] Among them, path parameters and query parameters are standard URL parameters. The data organization method carried in the path parameters and query parameters can be either of the two standard URL syntaxes that specify the attribute information of the first field to be transmitted in text. The data organization method carried in the encryption token can be to package and encrypt the attribute information of the first field to generate a ciphertext string and put it as a parameter in the path parameters and query parameters. The data organization method carried in the signature token can be to package and encrypt the attribute information and timeliness information of the first field and put the encryption result as a parameter in the path parameters and query parameters. Carrying encryption tokens and signature tokens can improve the security of the overall preview solution.
[0046] Using the above method, you can freely choose how to pass the attribute information of the first field, thereby improving the flexibility of the overall preview scheme.
[0047] In some cases, the request to obtain carries a first content range, which is smaller than the content range of the first multimodal content. The step of obtaining the first multimodal content based on the first physical address information and constructing a response to the request based on the obtained first multimodal content may include: mapping the first content range to a second content range relative to the first field; determining second physical address information matching the second content range based on the first physical address information; determining at least one fragmentation request based on the second physical address information, the fragmentation request being used to obtain third multimodal content, the third multimodal content corresponding to third physical address information, the third physical address information being determined based on the second physical address information, the second content range including all third multimodal content corresponding to the fragmentation requests; responding to the fragmentation request to obtain the third multimodal content; and constructing a response to the request based on the third multimodal content.
[0048] Understandably, in certain special scenarios (such as video dragging or file pagination loading), preview requests will automatically include a Range request header. This header indicates a first content range, which is smaller than the content range of the first multimodal content. In other words, a preview request with a Range request header indicates that a specific content range is being returned, not necessarily the entire file. This content range can be represented by the aforementioned start and end positions. The response header accompanying this preview request can include an Accept-Ranges header carrying bytes and a Content-Range header indicating the content range of the returned content segment. This content range includes a start and end position, and the Accept-Ranges header carrying bytes indicates support for streaming responses.
[0049] Furthermore, the first content range is the range of content to be previewed (i.e., the range in the preview request) determined from the client / browser's perspective. For example, the client / browser might describe wanting to preview content from 40MB (megabytes) to 60MB of the first multimodal content. Since the range corresponding to this 40MB to 60MB content in the first field is unknown, it is first necessary to determine the first content range within the first field. Therefore, the first content range can be mapped to a second content range relative to the first field, which reflects the range of the first content range within the first field.
[0050] Furthermore, the first physical address information describes the physical address of the complete first multimodal content. If it is necessary to locate the physical address of the first content range, it needs to be determined based on the first physical address information. Continuing with the example of the first physical address information above, if the first multimodal content is located under the associated object, with a starting offset of 150MB and a byte length of 100MB, and combined with the second content range (from the 40th to the 60th MB), it can be determined that the second physical address information matching the second content range is located under the associated object, with a starting offset of 190MB (150MB + 40MB) and an ending offset of 210MB (150MB + 60MB). Here, the starting offset and the starting position are equivalent concepts, and the ending offset and the ending position are equivalent concepts.
[0051] As an example, a fragmentation request can be determined based on the number of fragment files spanned by the second physical address information. For instance, if the second physical address information spans two fragment files, two fragmentation requests need to be generated. The third physical address information corresponding to each fragmentation request is obtained by dividing the data based on the second physical address information. It can be understood that the third multimodal content corresponding to adjacent fragmentation requests is continuous in content, and the third physical address information corresponding to adjacent fragmentation requests is also continuous.
[0052] The above method first determines the physical address information of the binary data of the multimodal content to be previewed. In addition, when fetching the binary data of the multimodal content, it can support fragmented fetching. Thus, when implementing the preview, the multimodal content can be continuously previewed based on the response corresponding to the continuous fragmented requests. This allows the client or browser to achieve progressive rendering, thereby solving the problems of high bandwidth consumption and long first-screen delay that are easy to occur in scenarios with large files (i.e., multimodal content) and batch preview.
[0053] In some cases, before obtaining the fourth multimodal content based on the fourth physical address information, the above-mentioned preview method for multimodal content in the AI data lake may further include the following steps: determining that access to the fourth physical address information is authorized; if the fourth physical address information includes the first physical address information, the fourth multimodal content includes the first multimodal content; if the fourth physical address information includes the third physical address information, the fourth multimodal content includes the third multimodal content.
[0054] By accessing the fourth physical address information, the corresponding fourth multimodal content can be obtained.
[0055] As an example, access to the fourth physical address information can be determined as authorized in at least one of the following ways: It is determined that the virtual uniform resource locator has a time limit; It is determined that the preview request carries a temporary credential, which is used to represent having permission to access the object.
[0056] By using the above method, after confirming that access to the physical address information is authorized, the corresponding multimodal content is read / fetched, thereby ensuring that binary data of multimodal content is read / fetched only within the authorized scope.
[0057] In some cases, target data can be obtained by: determining whether the target data exists in a cache during the process of obtaining the target data; and obtaining the target data from the cache in response to determining that the target data exists in the cache, wherein the target data includes at least one of the first metadata, the first physical address information, the first multimodal content, and the third multimodal content.
[0058] The first and third multimodal content here can be multimodal content with preview popularity values higher than preset popularity values, thus allowing caching to be configured in conjunction with preview popularity. Furthermore, the cache can contain binary data of the first multimodal content, and similarly, the cache can contain binary data of the third multimodal content.
[0059] A cache can be at least one of local memory cache, disk cache, and edge node cache. When there are multiple caches, the presence of the target data can be checked sequentially in each cache according to a priority strategy until all caches have been accessed. For example, the priority of each cache can be set according to access efficiency, with higher access efficiency resulting in higher priority. For instance, when the cache includes both local memory cache and disk cache, the presence of the target data can be checked first in the local memory cache. If the target data is not found in the local memory cache, the presence of the target data can be checked in the disk cache.
[0060] By prioritizing the retrieval of target data from the cache using the above methods, preview waiting time can be reduced, and issues such as repeated parsing and downloading can be resolved, thereby reducing preview latency.
[0061] In some cases, the cache can be set up according to key-value pairs, which include a query key and the corresponding target data. The query key is used to support the retrieval of the target data in the cache.
[0062] Depending on the target data, corresponding query keys can be constructed. For example, query keys can be constructed based on one or more of the following: object identifier (identifier of the multimodal dataset or identifier of the associated object of the multimodal dataset), field name of the first field, sample identifier, and content range. In this way, based on multiple types of information, query keys that are composite and correspond to the target data can be constructed from different granularities, improving the flexibility of the query.
[0063] In some cases, as can be seen from the above, the target data includes the first metadata. In this case, the preview method for multimodal content in the AI data lake may further include: in response to determining that the target data does not exist in the cache, obtaining dependent data; and obtaining the first metadata based on the dependent data.
[0064] The above attribute information can be used to construct a query key, and then the query for the first metadata can be performed in the cache based on the query key.
[0065] If it is determined from one or more of the aforementioned caches that the target data does not exist, then the dependency data can be obtained, which is used to support the acquisition of the first physical address information.
[0066] Using the above method, when the first metadata cannot be obtained from the cache, the first metadata is obtained based on the dependent data, so as to provide a data foundation for obtaining the first physical address information in the future.
[0067] In some cases, the dependent data includes a multimodal dataset, and in other cases, the dependent data includes a minimum metadata fragment, the size of which is smaller than the size of the multimodal dataset, and the minimum metadata fragment is used to characterize data sufficient to support obtaining the first physical address information.
[0068] As an example, the minimum metadata fragment may include multiple secondary metadata, which may include a version list, field descriptions, fragment file descriptions, column block indexes, page indexes, and offset and length information.
[0069] As can be seen from the above, multimodal datasets can support multiple version snapshots, lock the version list, filter irrelevant historical version fragments, and only load the index of the version currently being interacted with by the user.
[0070] The field description can record the field type of each column in the multimodal dataset, which is used to validate the first field and intercept invalid field preview requests in advance.
[0071] Multimodal datasets can be split into multiple fragment files (e.g., fragment_0.lance, fragment_1.lance, etc.). The fragment file description can record the range of line numbers (i.e. the range of sample identifiers) contained in the fragment file and the storage address of the fragment file. Therefore, based on the line number clicked by the user, it is possible to quickly locate which fragment file this data is in without traversing all fragments.
[0072] The column block index is used to record the storage location of the column corresponding to the first field in the fragment file.
[0073] Each fragment file is further divided into at least one data page. The page index can record the range of line numbers contained in each data page, which is used to locate the line number clicked by the user and which page of the fragment file it falls on, thus narrowing the search range.
[0074] Offset and length information: the offset is used to record the starting position of the first field in the fragment file, and the length information is used to record the byte length of the first field.
[0075] Using the methods described above, the first metadata used to obtain the first physical address information can be determined based on the information included in the minimum metadata fragment. Since the data size of the minimum metadata fragment is smaller than that of the multimodal dataset, downloading the entire multimodal dataset can be avoided, thus reducing resource consumption and preview latency.
[0076] In some cases, the above-mentioned method for previewing multimodal content in an AI data lake may further include the step of updating the cache in response to meeting preset conditions.
[0077] The preset conditions can include sub-conditions set in different dimensions. The first preset sub-condition is used to indicate that the preview heat value is lower than the preset heat value; The second preset sub-condition is used to characterize the version change of the multimodal dataset. The version change of the multimodal dataset may include changes in the metadata of the multimodal dataset (such as dataset identifier, first metadata, etc.). The third preset sub-condition is used to characterize that the cached data has expired. The validity of the cached data is determined by the set TTL (Time To Live). For example, the corresponding TTL can be set for the cached data when caching data.
[0078] By configuring preset conditions in the above manner, the cache can be updated based on the configured preset conditions, thus ensuring the reliability of the cache.
[0079] In some cases, with streaming, the next third-modal content to be rendered can be downloaded and cached in advance. When rendering the next third-modal content, it can be retrieved directly from the cache and reused, thereby reducing preview latency.
[0080] Figure 3 This is another schematic diagram illustrating the implementation environment of a preview method for multimodal content in an AI data lake, which may include a preview layer, an adaptation layer, a parsing layer, a planning layer, a response layer, and a caching layer. The preview layer supports user-triggered preview operations on the first field, generating a preview request to obtain the virtual unified resource locator (VURI) of the first field. The adaptation layer intercepts the preview request and constructs an acquisition request based on the attribute information of the VURI. The parsing layer parses the VURI in the acquisition request and constructs a query key based on the parsed attribute information. First metadata is obtained from the cache in the caching layer based on the query key. If the first metadata cannot be obtained directly from the cache, dependent data is obtained. Based on the dependent data, the first metadata is obtained. Based on the obtained first metadata, the mapping of the first physical address information and the execution scope is determined to obtain the second content scope. The planning layer determines at least one shard request based on the second physical address information. The response layer responds to the shard request to obtain the third multimodal content requested by the shard request and constructs a response to the acquisition request based on the third multimodal content. The preview layer, based on the response to the acquisition request, calls the rendering component to parse and render the binary data of the third multimodal content in the response, thereby achieving progressive rendering.
[0081] By using the above method, the preview request is converted into a request to obtain binary data for the multimodal content to be previewed. Then, the binary data is encapsulated to obtain the corresponding media resources that can be recognized by the browser's native rendering component (carried in the response), thereby achieving preview in a low-overhead manner.
[0082] Figure 4 This is a block diagram of a preview device for multimodal content in an AI data lake, shown in some scenarios, for reference. Figure 4The preview device 400 for multimodal content in the AI data lake includes: The first acquisition module 401 is used to obtain the virtual unified resource locator of the first field in response to the preview operation of the first field. The first field belongs to a field in the multimodal dataset. The first field is used to map the first multimodal content. The multimodal dataset is located in the AI data lake. The AI data lake presents the multimodal dataset in a first preset format. The first multimodal content belongs to any second multimodal content in the multimodal dataset. The virtual unified resource locator carries the attribute information of the first field. The first construction module 402 is used to construct an acquisition request based on the attribute information of the virtual uniform resource locator. The parsing module 403 is used to parse the virtual uniform resource locator in the request and obtain the first metadata based on the parsed attribute information; The second obtaining module 404 is used to obtain the first physical address information based on the first metadata; The second construction module 405 is used to obtain the first multimodal content based on the first physical address information, and to construct a response to the acquisition request based on the obtained first multimodal content, wherein the response is used to preview the first multimodal content.
[0083] In some cases, the first obtaining module 401 is used to: in response to the preview operation of the first field, construct a virtual Uniform Resource Locator (VURL) for the first field according to a second preset format to obtain the VURL for the first field. The second preset format includes multiple formats, and different data organization methods are used to encapsulate the attribute information of the first field into the VURL for different second preset formats.
[0084] In some cases, the request for obtaining content carries a first content range, which is smaller than the content range of the first multimodal content. The second construction module 405 is used to: Map the first content range to a second content range relative to the first field; Based on the first physical address information, determine the second physical address information that matches the second content range; Based on the second physical address information, at least one fragmentation request is determined. The fragmentation request is used to obtain third multimodal content. The third multimodal content corresponds to the third physical address information. The third physical address information is determined based on the second physical address information. The scope of the second content includes all the third multimodal content corresponding to the fragmentation requests. Respond to the fragmentation request to obtain the third multimodal content; Based on the third multimodal content, construct the response to the request.
[0085] In some cases, the preview device 400 for multimodal content in the AI data lake also includes: The determination module is used to determine whether access to the fourth physical address information is authorized before obtaining the fourth multimodal content based on the fourth physical address information; When the fourth physical address information includes the first physical address information, the fourth multimodal content includes the first multimodal content; when the fourth physical address information includes the third physical address information, the fourth multimodal content includes the third multimodal content. In some cases, target data is obtained in the following ways: During the process of obtaining the target data, it is determined whether the target data exists in the cache; In response to determining that the target data exists in the cache, the target data is obtained from the cache, the target data including at least one of the following: The first metadata; The first physical address information; The first multimodal content; Third, multimodal content.
[0086] In some cases, the target data includes the first metadata, and the preview device 400 for multimodal content in the AI data lake further includes: The acquisition module is configured to acquire dependent data in response to determining that the target data does not exist in the cache; The third acquisition module is used to acquire the first metadata based on the dependency data.
[0087] In some cases, the preview device 400 for multimodal content in the AI data lake also includes: An update module is used to update the cache in response to the fulfillment of preset conditions.
[0088] In some cases, the first physical address information includes the following: The object containing the first multimodal content includes the multimodal dataset or an associated object of the multimodal dataset; The starting offset of the first multimodal content; The byte length of the first multimodal content; The format type of the first multimodal content.
[0089] Among them, regarding Figure 4The implementation principles of each module in the preview device 400 for multimodal content in the AI data lake shown can be referred to the above content, and the preview device 400 for multimodal content in the AI data lake has the same technical effect as the above-mentioned management method for model service configuration.
[0090] Based on the same concept, a computer-readable medium is provided that stores a computer program thereon, wherein when executed by a processing device, the computer program causes the processing device to perform the steps of the above-described method, and the computer-readable medium has the technical effects that can be achieved by implementing the above-described related methods, and the technical effects can be referred to the above content.
[0091] Based on the same concept, a computer program product is provided, including a computer program, wherein when executed by a processor, the computer program causes the processor to implement the steps of the above-described method, and the computer program product has the technical effects that can be achieved by implementing the above-described related methods, and the technical effects can be referred to the above content.
[0092] Based on the same concept, an electronic device is provided, comprising: A storage device on which computer programs are stored; A processing device is used to execute the computer program in the storage device to implement the steps of the above method, and the electronic device has the technical effects that can be achieved by implementing the above-mentioned related methods, and the technical effects can be referred to the above content.
[0093] The following is for reference. Figure 5 It shows an electronic device suitable for implementing the above method (e.g. Figure 1 The diagram below shows the structure of the client (500). Terminal devices can include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs (Televisions), desktop computers, etc. Figure 5 The electronic device shown is merely an example and should not be construed as limiting its functionality or scope of use.
[0094] like Figure 5As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, the ROM 502, and the RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0095] Typically, the following devices can be connected to the input / output interface 505: input devices 506 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 507 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 508 including, for example, magnetic tape, hard disk, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0096] In particular, depending on certain circumstances, the processes described in the above-referenced flowchart can be implemented as computer software programs. For example, a computer program product is provided, comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. This computer program can be downloaded and installed from a network via communication device 509, or installed from storage device 508, or installed from read-only memory 502. When the computer program is executed by processing device 501, it performs the functions defined in the above-described methods.
[0097] It should be noted that the aforementioned computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM, or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In one case, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In another case, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.
[0098] In some scenarios, clients can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad-hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0099] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0100] The aforementioned computer-readable medium carries one or more programs. When the electronic device executes one or more of these programs, the electronic device causes the following: In response to a preview operation of a first field, it obtains a virtual Uniform Resource Locator (VURL) for the first field, wherein the first field belongs to a field in a multimodal dataset, the first field is used to map first multimodal content, the multimodal dataset is located in the AI data lake, the AI data lake presents the multimodal dataset in a first preset format, the first multimodal content belongs to any second multimodal content in the multimodal dataset, and the VURL carries attribute information of the first field; Based on the attribute information of the VURL, it constructs an acquisition request; Parses the VURL in the acquisition request and obtains first metadata based on the parsed attribute information; Obtains first physical address information based on the first metadata; Obtains the first multimodal content based on the first physical address information, and constructs a response to the acquisition request based on the obtained first multimodal content, the response being used to preview the first multimodal content.
[0101] Computer program code for performing the above operations can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages, as well as conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0102] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. Each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative cases, the functions indicated in the blocks may occur in a different order than those indicated in the figures. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0103] The modules mentioned above can be implemented in software or hardware. In some cases, the name of a module does not constitute a limitation on the module itself. For example, the first acquisition module can also be described as "in response to the preview operation of the first field, acquiring the virtual Uniform Resource Locator of the first field".
[0104] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field-Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application-Specific Standard Parts (ASSPs), Systems on Chips (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0105] In this context, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0106] The above description is merely illustrative and explains the technical principles employed. Those skilled in the art should understand that the scope of this document is not limited to the specific combinations of the above-described technical features, but should also cover any combination of the above-described technical features or their equivalents without departing from the above concept.
[0107] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. Multitasking and parallel processing may be advantageous in certain contexts. Similarly, while some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of this paper. Certain features described in the context of a single example can also be implemented in combination in a single example. Conversely, various features described in the context of a single example can also be implemented individually or in any suitable sub-combination in multiple examples.
[0108] Although this document has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims. Regarding the aforementioned apparatus, the specific manner in which the various modules perform their operations has already been described in detail in the section concerning the method, and will not be elaborated upon here.
Claims
1. A method for previewing multimodal content in an AI data lake, comprising: In response to the preview operation of the first field, a virtual uniform resource locator (VURL) for the first field is obtained. The first field belongs to a field in the multimodal dataset. The first field is used to map the first multimodal content. The multimodal dataset is located in the AI data lake. The AI data lake presents the multimodal dataset in a first preset format. The first multimodal content belongs to any second multimodal content in the multimodal dataset. The VURL carries the attribute information of the first field. Based on the attribute information of the virtual uniform resource locator, a request is constructed; Parse the Virtual Uniform Resource Locator in the request and obtain the first metadata based on the parsed attribute information; Based on the first metadata, obtain the first physical address information; Based on the first physical address information, the first multimodal content is obtained, and a response to the request is constructed based on the obtained first multimodal content. The response is used to preview the first multimodal content.
2. The method according to claim 1, wherein obtaining the virtual Uniform Resource Locator (URL) of the first field in response to a preview operation of the first field comprises: In response to the preview operation of the first field, a virtual Uniform Resource Locator (URL) for the first field is constructed according to a second preset format to obtain the URL for the first field. The second preset format includes multiple formats, and different data organization methods are used to encapsulate the attribute information of the first field into the URL.
3. The method according to claim 1, wherein the request for obtaining the content carries a first content range, the first content range being smaller than the content range of the first multimodal content, and the step of obtaining the first multimodal content based on the first physical address information and constructing a response to the request for obtaining the content based on the obtained first multimodal content includes: Map the first content range to a second content range relative to the first field; Based on the first physical address information, determine the second physical address information that matches the second content range; Based on the second physical address information, at least one fragmentation request is determined. The fragmentation request is used to obtain third multimodal content. The third multimodal content corresponds to the third physical address information. The third physical address information is determined based on the second physical address information. The scope of the second content includes all the third multimodal content corresponding to the fragmentation requests. Respond to the fragmentation request to obtain the third multimodal content; Based on the third multimodal content, construct the response to the request.
4. The method according to claim 1 or 3, before obtaining the fourth multimodal content based on the fourth physical address information, the method further includes: Access to the fourth physical address information was determined to be authorized; When the fourth physical address information includes the first physical address information, the fourth multimodal content includes the first multimodal content; when the fourth physical address information includes the third physical address information, the fourth multimodal content includes the third multimodal content.
5. The method according to claim 1 or 3, wherein the target data is obtained in the following manner: During the process of obtaining the target data, it is determined whether the target data exists in the cache; In response to determining that the target data exists in the cache, the target data is obtained from the cache, the target data including at least one of the following: The first metadata; The first physical address information; The first multimodal content; Third, multimodal content.
6. The method according to claim 5, wherein the target data includes the first metadata, and the method further includes: In response to the determination that the target data does not exist in the cache, the dependent data is retrieved; Based on the dependency data, the first metadata is obtained.
7. The method according to claim 5, further comprising: The cache is updated in response to the fulfillment of preset conditions.
8. The method according to claim 1, wherein the first physical address information includes the following information: The object containing the first multimodal content includes the multimodal dataset or an associated object of the multimodal dataset; The starting offset of the first multimodal content; The byte length of the first multimodal content; The format type of the first multimodal content.
9. A preview device for multimodal content in an AI data lake, comprising: The first acquisition module is used to obtain the virtual unified resource locator of the first field in response to the preview operation of the first field. The first field belongs to a field in the multimodal dataset. The first field is used to map the first multimodal content. The multimodal dataset is located in the AI data lake. The AI data lake presents the multimodal dataset in a first preset format. The first multimodal content belongs to any second multimodal content in the multimodal dataset. The virtual unified resource locator carries the attribute information of the first field. The first construction module is used to construct an acquisition request based on the attribute information of the virtual uniform resource locator; The parsing module is used to parse the Virtual Uniform Resource Locator in the request and obtain the first metadata based on the parsed attribute information; The second obtaining module is used to obtain the first physical address information based on the first metadata; The second construction module is used to obtain the first multimodal content based on the first physical address information, and to construct a response to the acquisition request based on the obtained first multimodal content, wherein the response is used to preview the first multimodal content.
10. An electronic device, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-8.