Method and device for accessing heterogeneous data
Through the high-order tensor spatial model, the problem of integrated big data in vehicle-road and cloud-based big data access is solved, and the rapid access and efficient analysis of data is realized, and a concise and efficient theoretical model is provided.
Patent Information
- Application Number
- CN202411576067.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-06
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2044-11-06
AI Technical Summary
The existing technology is difficult to meet the needs of integrated big data access for vehicle-road and cloud, especially the aggregation and fusion of multi-source heterogeneous data, which has problems such as large data structure differences, many sources and widespread, high value density, and real-time updates.
High-order tensor spatial model is used to uniformly represent multi-source heterogeneous data, and fast access and efficient analysis of data are achieved by obtaining low-order sub-tensors, matrix singular value decomposition and modular multiplication processing.
A unified representation of unstructured, semi-structured and structured data is realized. A simple and efficient theoretical model provides the basis for integrated big data analysis and processing of vehicles and roads, reduces data complexity, and obtains high-quality data sets.
Smart Images

Figure CN119513574B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the fields of smart transportation and big data, and in particular to a method and device for accessing heterogeneous data. Background Art
[0002] Vehicle-road-cloud integration is a key technological development direction for the future convergence of urban transportation towards digitalization. This integration encompasses vehicles on the road, people walking along the roadside, various roadside equipment, and cloud-based dispatching and information monitoring systems. It relies on roadside sensing, edge computing, cloud-based information fusion, and key technologies such as C-V2X (Cellular Vehicle-to-Everything) and 4G (4th generation mobile communication technology) or 5G (5th generation mobile communication technology) to achieve comprehensive coordination and collaboration between the vehicle, road, and cloud, including collaborative perception, control, and decision-making and planning. Based on this technological foundation, cities can form a "traffic brain" to achieve unified digital management of urban traffic conditions. These are the current development directions for the convergence of urban transportation towards digitalization.
[0003] However, due to the large structural differences in vehicle-road-cloud data, the large number and breadth of data sources, the high value density of data, and real-time data updates, existing big data processing technologies are difficult to meet the needs of integrated vehicle-road-cloud big data access. Summary of the Invention
[0004] The purpose of this application is to solve at least one of the above technical deficiencies.
[0005] In one aspect, an embodiment of the present application provides a method for accessing heterogeneous data, the method comprising:
[0006] Obtain a high-order tensor space model. This model includes a unified representation of multi-source heterogeneous data in vehicle-road-cloud integration. The high-order tensor space model is obtained by fusing low-order sub-tensors of each multi-source heterogeneous data into a tensor base space model. The order of each low-order sub-tensor is determined based on the attribute information of the corresponding multi-source heterogeneous data.
[0007] Determining an order tensor model of data to be extracted from a high-order tensor space model;
[0008] Perform modular expansion on the order of the order tensor model to obtain the corresponding matrix;
[0009] Perform matrix singular value decomposition on the matrix to obtain the corresponding decomposition result;
[0010] Perform modular multiplication on the decomposition result to obtain the data to be extracted.
[0011] Optionally, the multi-source heterogeneous data includes at least one of real-time message data in vehicle-road-cloud integration, various structured report data, attribute data, unstructured text and image data, and various video and voice streaming data.
[0012] Optionally, the unified representation of the multi-source heterogeneous data is obtained by the following method:
[0013] Obtaining various multi-source heterogeneous data and attribute information corresponding to each multi-source heterogeneous data;
[0014] According to the attribute information corresponding to each multi-source heterogeneous data, the low-order sub-tensor corresponding to each multi-source heterogeneous data is determined, and different orders of the low-order sub-tensor represent different attribute information;
[0015] Obtaining a tensor basis space model, the tensor basis space model including at least three orders corresponding to basic attribute information;
[0016] Arrange each multi-source heterogeneous data into the tensor basis space model according to the order of the low-order sub-tensor corresponding to each multi-source heterogeneous data and the order corresponding to each basic attribute information included in the tensor basis space model to obtain a high-order tensor space model;
[0017] A unified representation of multi-source heterogeneous data is obtained based on a high-order tensor space model.
[0018] Optionally, the basic attribute information includes time information, space information, and user information;
[0019] The high-order tensor space model is obtained by the following formula:
[0020] X=T u ∪T semi ∪T s =(time,space,user,i1,…,i N ,V)
[0021] Among them, X represents the P+3 order tensor, time represents time information, space represents space information, user represents user information, i1,…,i N Indicates the order of the object that can be expanded, and V indicates the value.
[0022] Optionally, perform matrix singular value decomposition on the matrix using the following formula to obtain the corresponding decomposition result:
[0023]
[0024] Among them, × k ,k=1,2,…,N represents the modal product, is the order tensor model of the data to be extracted; A (1) ,A (2) ,…,A (N) are factor matrices of different dimensions, is the core tensor.
[0025] Optionally, if the data to be extracted is special data, and the special data is at least one of data with incomplete attribute information, data with inconsistent attribute information, redundant data, and noise data in the multi-source heterogeneous data, the method further includes:
[0026] Determine the rank tensor model of special data from the higher-rank tensor space model;
[0027] Perform multi-level high-order singular value decomposition on the order tensor model of special data to obtain the data corresponding to the special data.
[0028] Optionally, perform multi-level high-order singular value decomposition on the order tensor model of special data using the following formula:
[0029] X=U1U2…U m W m +U s W s
[0030] Among them, X is the order tensor model, U1U2…U m are basis matrices of different orders, W s ,W m is the corresponding coefficient matrix.
[0031] Optionally, the method further includes:
[0032] Get the newly added data and determine the low-order sub-tensor corresponding to the newly added data;
[0033] If the high-order tensor space model does not include the order of the low-order sub-tensor corresponding to the newly added data, the excluded order is added to the high-order tensor space model as an expandable object order to obtain an expanded high-order tensor space model, and the newly added data is uniformly represented based on the expanded high-order tensor space model;
[0034] If the high-order tensor space model includes the order of the low-order sub-tensors corresponding to the newly added data, the same order is merged and the different orders are retained in the high-order tensor space model to obtain a unified representation of the newly added data.
[0035] Optionally, each order in the high-order tensor space model includes at least one dimension, and the same order is merged in the high-order tensor space model, including:
[0036] In high-order tensor space models, the same order is merged using the finest-grained dimension.
[0037] On the other hand, an embodiment of the present application provides a device for accessing heterogeneous data, including:
[0038] The model acquisition module is used to obtain a high-order tensor space model. The high-order tensor space model includes a unified representation of multi-source heterogeneous data in vehicle-road-cloud integration. The high-order tensor space model is obtained by fusing low-order sub-tensors of each multi-source heterogeneous data into a tensor base space model. The order of each low-order sub-tensor is determined according to the attribute information of the corresponding multi-source heterogeneous data.
[0039] A model determination module, used for determining the order tensor model of the data to be extracted from the high-order tensor space model;
[0040] The data extraction module is used to perform modular expansion on the order of the order tensor model to obtain the corresponding matrix, perform matrix singular value decomposition on the matrix to obtain the corresponding decomposition result, and perform modular multiplication on the decomposition result to obtain the data to be extracted.
[0041] Optionally, the multi-source heterogeneous data includes at least one of real-time message data in vehicle-road-cloud integration, various structured report data, attribute data, unstructured text and image data, and various video and voice streaming data.
[0042] Optionally, the unified representation of the multi-source heterogeneous data is obtained by the following method:
[0043] Obtaining various multi-source heterogeneous data and attribute information corresponding to each multi-source heterogeneous data;
[0044] According to the attribute information corresponding to each multi-source heterogeneous data, the low-order sub-tensor corresponding to each multi-source heterogeneous data is determined, and different orders of the low-order sub-tensor represent different attribute information;
[0045] Obtaining a tensor basis space model, the tensor basis space model including at least three orders corresponding to basic attribute information;
[0046] Arrange each multi-source heterogeneous data into the tensor basis space model according to the order of the low-order sub-tensor corresponding to each multi-source heterogeneous data and the order corresponding to each basic attribute information included in the tensor basis space model to obtain a high-order tensor space model;
[0047] A unified representation of multi-source heterogeneous data is obtained based on a high-order tensor space model.
[0048] Optionally, the basic attribute information includes time information, space information, and user information;
[0049] The high-order tensor space model is obtained by the following formula:
[0050] X=T u ∪T semi ∪T s =(time,space,user,i1,…,i N ,V)
[0051] Among them, X represents the P+3 order tensor, time represents time information, space represents space information, user represents user information, i1,…,i N Indicates the order of the object that can be expanded, and V indicates the value.
[0052] Optionally, the data extraction module performs matrix singular value decomposition on the matrix using the following formula to obtain a corresponding decomposition result:
[0053]
[0054] Among them, × k ,k=1,2,…,N represents the modal product, is the order tensor model of the data to be extracted; A (1) ,A (2) ,…,A (N) are factor matrices of different dimensions, is the core tensor.
[0055] Optionally, if the data to be extracted is special data, which is at least one of data with incomplete attribute information, data with inconsistent attribute information, redundant data, and noise data in the multi-source heterogeneous data, the data extraction module is further configured to:
[0056] Determine the rank tensor model of special data from the higher-rank tensor space model;
[0057] Perform multi-level high-order singular value decomposition on the order tensor model of special data to obtain the data corresponding to the special data.
[0058] Optionally, the data extraction module performs multi-level high-order singular value decomposition on the order tensor model of special data using the following formula:
[0059] X=U1U2…U m W m +U s W s
[0060] Among them, X is the order tensor model, U1U2…U m are basis matrices of different orders, W s ,W m is the corresponding coefficient matrix.
[0061] Optionally, the device further includes an expansion module, specifically configured to:
[0062] Get the newly added data and determine the low-order sub-tensor corresponding to the newly added data;
[0063] If the high-order tensor space model does not include the order of the low-order sub-tensor corresponding to the newly added data, the excluded order is added to the high-order tensor space model as an expandable object order to obtain an expanded high-order tensor space model, and the newly added data is uniformly represented based on the expanded high-order tensor space model;
[0064] If the high-order tensor space model includes the order of the low-order sub-tensors corresponding to the newly added data, the same order is merged and the different orders are retained in the high-order tensor space model to obtain a unified representation of the newly added data.
[0065] Optionally, each order in the high-order tensor space model includes at least one dimension, and the extension module merges the same orders in the high-order tensor space model, specifically for:
[0066] In high-order tensor space models, the same order is merged using the finest-grained dimension.
[0067] In another aspect, an embodiment of the present application provides an electronic device, including a processor and a memory:
[0068] The memory is configured to store machine-readable instructions, which, when executed by the processor, cause the processor to perform any one of the methods for accessing heterogeneous data.
[0069] The beneficial effects of the technical solutions provided in the embodiments of the present application include at least:
[0070] In the embodiment of the present application, a high-order tensor space unified representation model is proposed based on the characteristics of the attribute information of data with different structures. On the basis of ensuring the completeness of the original data characteristics, it not only realizes the unified representation of unstructured data, semi-structured data, and structured data (i.e., tensor model), but also establishes a concise and efficient theoretical representation model for the analysis and processing of vehicle-road-cloud integrated big data. At the same time, when accessing data later, the unified representation model corresponding to the data to be accessed can be quickly determined based on the attribute information of different heterogeneous data, and then the unified representation model is subjected to singular value decomposition to obtain the data to be accessed. At this time, not only can the data be quickly accessed, but the obtained data is more suitable for subsequent big data analysis and mining. In addition, in the embodiment of the present application, the finest granularity dimension can be used for merging at the same order. At this time, a core high-quality data set can be obtained, and the order of this data set is the same as that of the unified tensor representation model, but the dimension will be smaller, reducing the complexity of the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0072] Figure 1 A schematic diagram of a process for accessing heterogeneous data provided in an embodiment of the present application;
[0073] Figure 2 A schematic diagram of the internal structure of a tensor basis space model provided in an embodiment of the present application;
[0074] Figure 3 A schematic diagram of the structure of the vehicle-road-cloud integrated big data aggregation and fusion platform provided in an embodiment of the present application;
[0075] Figure 4 A schematic diagram of a data storage architecture provided in an embodiment of the present application;
[0076] Figure 5 A schematic diagram of the structure of a device for accessing heterogeneous data provided in an embodiment of the present application;
[0077] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0078] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present invention.
[0079] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.
[0080] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0081] Vehicle-road-cloud integration is an important technical development direction for the future integration of urban transportation towards "digital". In the past year, a large number of policies have been issued to promote the implementation of vehicle-road-cloud integration. However, since vehicle-road-cloud integration involves multiple departments, the aggregation, sharing and integration of data between departments are all centered on their respective applications, and a complex network structure is formed between multiple departments and systems, which poses new challenges to the production, use, maintenance and management of data. For example, there are great difficulties in the aggregation and integration of multi-source heterogeneous data such as real-time message data, various structured report data and attribute data, unstructured text and pictures, various video and voice streaming data, and how to access multi-source heterogeneous data. Based on this, the present application provides a method for accessing heterogeneous data and an electronic device, which aims to solve the above technical problems of the prior art.
[0082] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0083] Specifically, such as Figure 1 As shown, the method may include:
[0084] Step S101, obtain a high-order tensor space model. The high-order tensor space model includes a unified representation of various multi-source heterogeneous data in the vehicle-road-cloud integration. The high-order tensor space model is obtained by fusing the low-order sub-tensors of various multi-source heterogeneous data into the tensor basis space model. The order of each low-order sub-tensor is determined according to the attribute information of the corresponding multi-source heterogeneous data.
[0085] Optionally, the multi-source heterogeneous data includes at least one of real-time message data in vehicle-road-cloud integration, various structured report data, attribute data, unstructured text and image data, and various video and voice streaming data.
[0086] Among them, vehicle-road-cloud integration refers to the integration of vehicles, roads and clouds in urban transportation. Multi-source heterogeneous data refers to data from different sources. In the embodiment of the present application, the multi-source heterogeneous data may refer to real-time message data, various structured report data, attribute data, unstructured text and image data, various video and voice streaming data, etc. in the vehicle-road-cloud integration, and the embodiment of the present application does not limit this. The high-order tensor space model is obtained by fusing the low-order sub-tensors of each multi-source heterogeneous data into the tensor basis space model, and the order of each low-order sub-tensor is determined according to the attribute information of the corresponding multi-source heterogeneous data.
[0087] Optionally, the attributes of different heterogeneous data may be the same or different, and this is not limited in the present embodiment. For example, for unstructured video data, its attribute information may include video frame, image width, image height, color space, etc., while the attribute information of structured XML (Extensible Markup Language) document data may include row and column, ASCII (American Standard Code for Information Interchange) encoding of elements, etc.
[0088] In an optional embodiment of the present application, the unified representation of the multi-source heterogeneous data is obtained by the following method:
[0089] Obtaining various multi-source heterogeneous data and attribute information corresponding to each multi-source heterogeneous data;
[0090] According to the attribute information corresponding to each multi-source heterogeneous data, the low-order sub-tensor corresponding to each multi-source heterogeneous data is determined, and different orders of the low-order sub-tensor represent different attribute information;
[0091] Obtaining a tensor basis space model, the tensor basis space model including at least three orders corresponding to basic attribute information;
[0092] Arrange each multi-source heterogeneous data into the tensor basis space model according to the order of the low-order sub-tensor corresponding to each multi-source heterogeneous data and the order corresponding to each basic attribute information included in the tensor basis space model to obtain a high-order tensor space model;
[0093] A unified representation of multi-source heterogeneous data is obtained based on a high-order tensor space model.
[0094] Optionally, when acquiring each multi-source heterogeneous data set, attribute information corresponding to each multi-source heterogeneous data set can be obtained. For each multi-source heterogeneous data set, the corresponding attribute information can characterize the multi-source heterogeneous data set. Furthermore, based on the attribute information corresponding to each multi-source heterogeneous data set, a low-order sub-tensor corresponding to each multi-source heterogeneous data set can be determined. In this case, for each multi-source heterogeneous data set, each order in the low-order sub-tensor determined for the multi-source heterogeneous data set represents a different attribute information.
[0095] For different heterogeneous data, the corresponding low-order sub-tensors are determined in different ways. For unstructured video data, a video data in MPEG4 (Moving Pictures Experts Group) format can be represented as a fourth-order sub-tensor:
[0096]
[0097] Among them, I f , I w , I h , I c Represent the video frame, image width, image height, and color space of the video data respectively, and the element value of the fourth-order tensor is equal to the encoding value of the video. For example, a video clip with 750 frames, a resolution of 768×576, and MPEG-4 format can be represented as the following fourth-order tensor:
[0098]
[0099] The color space is represented by three colors: red, green, and blue.
[0100] Optionally, for semi-structured XML document data, when determining its low-order sub-tensor, since the XML document has a hierarchical structure, it can be parsed into a tree structure. Therefore, the XML document can be represented as a third-order sub-tensor:
[0101]
[0102] Among them, I er ,I ec ,I en They represent the ASCII codes of the rows, columns and elements of the identity matrix respectively.
[0103] Alternatively, for structured database tables, since relational databases often store structured data, and in simple database tables, fields are often represented by numbers or characters, they can be represented as a matrix. For complex fields, such as BLOB (binary large object) types, new tensor rank can be added to represent them.
[0104] Correspondingly, after the unstructured, semi-structured, and structured data contained in the multi-source heterogeneous data are converted into corresponding low-order sub-tensors, the tensor basis space model can be obtained. Since the tensor basis space model includes at least three orders corresponding to basic attribute information, such as time I t 、Space I s User I u Three basic attribute information, at this time, according to the order of the low-order sub-tensor corresponding to each multi-source heterogeneous data and the order corresponding to each basic attribute information included in the tensor basis space model, each multi-source heterogeneous data can be arranged into the tensor basis space model to obtain a high-order tensor space model. Specifically, Figure 2 The internal structure of the tensor basis space model shown in the figure shows that the tensor basis space is a double-layer space and is a third-order (i.e., I t , I s and I u ) spatial representation model, after encoding multi-source heterogeneous data (i.e., pictures, trajectories, videos, and XML documents in the figure), they can be uniformly incorporated into the tensor-based space model. At this time, the identity and structure of the original data (i.e., each multi-source heterogeneous data) will be losslessly preserved in the tensor during the representation stage, which means that a unified representation of each multi-source heterogeneous data is obtained.
[0105] Optionally, the process of arranging the multi-source heterogeneous data into the tensor basis space model can be expressed by the following formula:
[0106] f:(d u ∪d semi ∪d s )→T u ∪T semi ∪T s
[0107] Among them, the independent variables represent unstructured data d u , semi-structured data semi and structured data d s , T u ,T semi ,T s They represent the low-order sub-tensors (a second-order matrix) corresponding to the corresponding unstructured data, semi-structured data, and structured data, respectively, and ∪ represents the union.
[0108] In the embodiment of the present application, the basic attribute information may include time information, space information, and user information, and the high-order tensor space model may be obtained by the following formula:
[0109] X=T u ∪T semi ∪T s =(time,space,user,i1,…,i N ,V)
[0110] Among them, X represents the P+3 order tensor, time represents time information, space represents space information, user represents user information, i1,…,i N Indicates the order of the object that can be expanded, and V indicates the value.
[0111] At this time, after obtaining the high-order tensor space model, each multi-source heterogeneous data obtains a corresponding unified variable data structure in the computer system (that is, uniformly represented as a P+3 order tensor), and through the above processing, large-scale heterogeneous multi-source data is converted into a unified two-layer space.
[0112] Step S102 , determining the 1-order tensor model of the data to be extracted from the high-order tensor space model, and performing modular expansion on the 1-order tensor model to obtain a corresponding matrix.
[0113] Step S103: perform matrix singular value decomposition on the matrix to obtain a corresponding decomposition result.
[0114] Step S104: performing modular multiplication processing on the decomposition result to obtain the data to be extracted.
[0115] Optionally, after all multi-source heterogeneous data are uniformly represented as a tensor model (i.e., a P+3-order tensor), if data extraction is required, the tensor model of the data to be extracted can be determined from the high-order tensor space model. The order of the tensor model is then modularly expanded to obtain the corresponding matrix. Each order reflects the internal structure of the heterogeneous data from a different perspective. For example, for a P-order tensor model, the P matrices obtained after modular expansion represent the characteristics of the heterogeneous data at each tensor order from P perspectives.
[0116] Furthermore, the obtained matrix can be subjected to matrix singular value decomposition, and the decomposition result can be subjected to modular multiplication, at which point the core data of the data to be extracted can be extracted.
[0117] In an optional embodiment of the present application, the matrix can be subjected to matrix singular value decomposition using the following formula to obtain the corresponding decomposition result:
[0118]
[0119] Among them, × k ,k=1,2,…,N represents the modal product, is the order tensor model of the data to be extracted; A (1) ,A (2) ,…,A (N) are factor matrices of different dimensions, is the core tensor.
[0120] Optional, A (1) ,A (2) ,…,A (N) It is usually considered as the principal components in different dimensions, and Each element represents the degree of interaction between different components. The formula can be used to obtain the high-order singular value decomposition (HOSVD) and A (1) ,A (2) ,…,A (N) , now there is and A (1) ,A (2) ,…,A (N) , we can calculate the high-order tensor by the following formula Features in each dimension:
[0121]
[0122] in, is a tensor The n-th order feature of is the core tensor The matrix obtained by expanding the nth dimension is, represents the Kronecker product.
[0123] In an optional embodiment of the present application, if the data to be extracted is special data, and the special data is at least one of data with incomplete attribute information, data with inconsistent attribute information, redundant data, and noise data in the multi-source heterogeneous data, the method further includes:
[0124] Determine the rank tensor model of special data from the higher-rank tensor space model;
[0125] Perform multi-level high-order singular value decomposition on the order tensor model of special data to obtain the data to be extracted.
[0126] In practice, vehicle-road-cloud integration also includes some special data. This special data may be at least one of data with incomplete attribute information, data with inconsistent attribute information, redundant data, and noisy data. This embodiment of the present application does not limit this. Optionally, if the data to be extracted is special data, after determining the order tensor model of the special data from the high-order tensor space model, the order tensor model of the special data can be subjected to multi-level high-order singular value decomposition to obtain the data to be extracted. The data obtained in this case is more suitable for subsequent big data analysis and mining.
[0127] In an optional embodiment of the present application, a multi-level high-order singular value decomposition can be performed on the order tensor model of special data using the following formula:
[0128] X=U1U2…U m W m +U s W s
[0129] Among them, X is the order tensor model, U1U2…U m are basis matrices of different orders, W s ,W m is the corresponding coefficient matrix.
[0130] Optionally, performing multi-level high-order singular value decomposition on the order tensor model can be interpreted as performing multi-level decomposition on the attribute information of the high-order tensor X to find its new representation W m ,W s and the new representation space D1D2…D m ,D s ,Right now:
[0131]
[0132] In practical applications, in order to increase the nonlinear representation capability of the basis space, the coefficient matrix can be modified by the following formula according to the deep neural network method:
[0133] W i =g(D i+1 W i+1 )
[0134] Among them, g(D i+1 W i+1 ) is a nonlinear activation function. Therefore, the following objective function can be constructed:
[0135]
[0136] Among them, ‖·‖2 represents the 2-norm. The objective function of the above formula is solved using a solution method similar to the stacked autoencoder network, which is divided into two stages: layer-by-layer pre-solution and overall fine-tuning:
[0137] (1) The pre-solution stage specifically includes two processes, A and B:
[0138] A. Let X = U1W1 + U s W s , solve the minimization problem: Complete the first level of decomposition;
[0139] B. Continue to decompose W1 into W1=U2W2 to complete the second level of decomposition;
[0140] By looping in this way, all layers can be pre-solved, and then by greedy decomposition layer by layer, the solution of each layer becomes a traditional optimization learning problem.
[0141] (2) Overall fine-tuning stage:
[0142] In practical applications, in order to make the solution more accurate, after the pre-solution stage, the solution can be optimized by minimizing the loss function and the stochastic gradient descent method (that is, entering the overall fine-tuning stage).
[0143] In an optional embodiment of the present application, the method further includes:
[0144] Get the newly added data and determine the low-order sub-tensor corresponding to the newly added data;
[0145] If the high-order tensor space model does not include the order of the low-order sub-tensor corresponding to the newly added data, the excluded order is added to the high-order tensor space model as an expandable object order to obtain an expanded high-order tensor space model, and the newly added data is uniformly represented based on the expanded high-order tensor space model;
[0146] If the high-order tensor space model includes the order of the low-order sub-tensors corresponding to the newly added data, the same order is merged and the different orders are retained in the high-order tensor space model to obtain a unified representation of the newly added data.
[0147] In practice, in the field of vehicle-road-cloud integration, the data will be constantly changed and updated. At this time, for the newly added data, we can first determine the low-order sub-tensor corresponding to the newly added data, and then compare the order of the low-order sub-tensor with the order in the high-order tensor space model. If the high-order tensor space model does not include the order of the low-order sub-tensor corresponding to the newly added data, the excluded order will be added to the high-order tensor space model as an expandable object order to obtain the expanded high-order tensor space model, and then the newly added data will be uniformly represented based on the expanded high-order tensor space model. On the contrary, if the high-order tensor space model includes the order of the low-order sub-tensor corresponding to the newly added data, the same order can be merged in the high-order tensor space model, and the different orders can be retained to obtain a unified representation of the newly added data.
[0148] For example, suppose and Representing two fourth-order tensors, the tensor expansion method is as follows:
[0149]
[0150] In the above formula, C is the result of the expansion, and the expansion operator Satisfies the associative law, that is The order of the sub-tensor can be extended to the order of the existing high-order tensor space model in different directions. The tensor expansion operator merges the same order and retains the different orders.
[0151] In an embodiment of the present application, each order in the high-order tensor space model includes at least one dimension, and the same orders are merged in the high-order tensor space model, including:
[0152] In high-order tensor space models, the same order is merged using the finest-grained dimension.
[0153] In practical applications, each order of the sub-tensor has many dimensions. When merging sub-tensors of the same order, the finest-grained dimensions can be used for merging at the same order, that is, through dimensionality reduction. At this time, a core high-quality data set can be obtained, and the order of this data set is the same as the unified tensor representation model, but the dimension will be smaller, reducing the complexity of the data.
[0154] For example, two subtensors T sub1 and T sub2 Both contain time steps, denoted as I t-1 and I t-2 , the dimensions of these two orders are I t-1 ∈{i1,i2} and I t-2 ∈{i1,i3}, the new tensor model constructed using the tensor expansion operator is expressed as The dimension of the time order is I t ∈{i1,i2,i3}.
[0155] In this application's embodiments, a unified high-order tensor space representation model is proposed to address the characteristics of data with different structures. This model achieves unified representation for unstructured, semi-structured, and structured data, establishing a concise and efficient theoretical representation model for integrated vehicle-road-cloud big data analysis and processing. Furthermore, to address the issue of feature conflicts in heterogeneous data, a dynamic tensor space fusion mechanism is proposed. This achieves efficient, unified representation of heterogeneous data in a high-order space while ensuring the completeness of the original data's features.
[0156] The embodiment of the present application provides a vehicle-road-cloud integrated big data aggregation and fusion platform for executing a method for accessing heterogeneous data in the above-mentioned embodiment of the application. Among them, the vehicle-road-cloud integrated big data aggregation and fusion platform adopts Hadoop (Hadoop Distributed File System) distributed system and HDFS (Hadoop Distributed File System) distributed storage. It internally integrates JDBC (Java Database Connectivity, defines application program interface), ODBC (Open Database Connectivity, open database connection), Kafka (open source stream processing platform), and Sqoop (conversion tool) components to seamlessly connect data from traditional relational databases to HDFS. The computing layer adopts Apache HBase (Hadoop Database, non-relational distributed database) real-time online data processing and Hive (data warehouse tool) computing execution engine. Specifically, the structure of the vehicle-road-cloud integrated big data aggregation and fusion platform is as follows: Figure 3 As shown, it mainly includes five layers: multi-source heterogeneous data aggregation layer, data quality management layer, big data storage layer, data security sharing layer and data service layer; two guarantee systems: security assurance system and standardization system.
[0157] Among them, the multi-source heterogeneous data aggregation layer is oriented towards data of types such as people, vehicles, and roads, targeting real-time message data (such as vehicle trajectory, personnel positioning), various structured report data (such as city data, business data), attribute data (such as population data, event data), unstructured text and images (such as text data, image data), various types of video and voice streaming data (such as key areas of traffic checkpoints) and other data (such as GIS data, department offline), and multi-source heterogeneous data aggregation is carried out according to real-time requirements, and each aggregation link dynamically loads and balances the load according to the load situation.
[0158] The data quality management layer is mainly responsible for data extraction, data cleaning, data collection, data dimensionality reduction (not shown in the figure), data conversion, data association, data comparison and data verification.
[0159] The big data storage layer uses the Hadoop distributed memory list storage system (i.e. HDFS, HBase and Hive in the figure) for data storage. HDFS uses a three-copy strategy to ensure data security and reliability. The data storage architecture is as follows: Figure 4 As shown, it includes distributed memory list storage, parallel algorithm library, data mining (i.e. DataMning in the figure), batch computing (MapReduce), distributed online database Hyperbase, unified resource scheduling YARN, Sparkd management and control, distributed file system HDFS and redundant coding Erasure Code. Among them, the distributed NOSQL (Not Only SQ, non-relational database) online database Hyperbase (a database) is provided on top of HDFS to provide platform support for high-concurrency retrieval analysis and transaction support. Hyperbase supports multi-dimensional millisecond-level global indexing, full-text indexing, combined indexing and other retrieval queries for massive data through multiple indexes. The platform's big data storage layer supports low-cost storage of various structured, semi-structured, and unstructured massive data, providing basic support for the storage and use of massive historical data. Hyperbase provides high-concurrency and low-latency retrieval capabilities and provides high-performance data access services to the outside world.
[0160] At the data security sharing level, in response to the sharing needs between different departments, different applications and different businesses, the platform opens different permissions to ensure unified allocation of resources and permission management based on data requirements such as offline / streaming data type, data unit KB (Kilobyte, kilobyte) / MB (Mbyte, megabyte) / GB (Gigabyte, gigabyte) / TB (Terabyte, billion bytes), data real-time requirements weekly / monthly / real-time, data security level requirements, and data encryption requirements. The platform includes data query, data upload, data synchronization (not shown in the figure), data download, data analysis, data templates, application access authorization, commission and bureau management, user management, role management, log monitoring, data review, data directory, data security management, etc.
[0161] The data service layer mainly provides services such as search engines, workflow engines, Restful (the design style and development method of network applications) service engines and knowledge fusion engines. Specifically, it can include Feemarker (a template engine), J2EE (Java 2Platform, Enterprise Edition, an enterprise-level Java application development platform), XML (Xtensible Markup Language, a markup language), Json (JavaScript Object Notation, a lightweight data exchange format), Token (string), HITP (Hypertext Transfer Protocol, a stateless, request / response model-based protocol), etc.
[0162] The embodiments of the present application aim to solve the problem of aggregation and sharing of multi-source heterogeneous data, such as real-time message data, structured report data and attribute data, unstructured text and images, and video and voice streaming data, and propose to build a unified big data aggregation, fusion and sharing platform with the ability to process large-scale heterogeneous data sources, so as to fundamentally and globally solve the long-standing problems of "independent management, fragmentation, numerous silos and information islands" in informatization construction.
[0163] The embodiment of the present application provides a device for accessing heterogeneous data, such as Figure 5 As shown, the access device 50 may include: a model acquisition module 501, a model determination module 502 and a data extraction module 503, wherein:
[0164] Model acquisition module 501 is used to obtain a high-order tensor space model. The high-order tensor space model includes a unified representation of multi-source heterogeneous data in vehicle-road-cloud integration. The high-order tensor space model is obtained by fusing low-order sub-tensors of each multi-source heterogeneous data into a tensor base space model. The order of each low-order sub-tensor is determined based on the attribute information of the corresponding multi-source heterogeneous data.
[0165] A model determination module 502 is used to determine the order tensor model of the data to be extracted from the high-order tensor space model;
[0166] The data extraction module 503 is used to perform modular expansion on the order of the order tensor model to obtain the corresponding matrix, perform matrix singular value decomposition on the matrix to obtain the corresponding decomposition result, and perform modular multiplication processing on the decomposition result to obtain the data to be extracted.
[0167] Optionally, the multi-source heterogeneous data includes at least one of real-time message data in vehicle-road-cloud integration, various structured report data, attribute data, unstructured text and image data, and various video and voice streaming data.
[0168] Optionally, the unified representation of the multi-source heterogeneous data is obtained by the following method:
[0169] Obtaining various multi-source heterogeneous data and attribute information corresponding to each multi-source heterogeneous data;
[0170] According to the attribute information corresponding to each multi-source heterogeneous data, the low-order sub-tensor corresponding to each multi-source heterogeneous data is determined, and different orders of the low-order sub-tensor represent different attribute information;
[0171] Obtaining a tensor basis space model, the tensor basis space model including at least three orders corresponding to basic attribute information;
[0172] Arrange each multi-source heterogeneous data into the tensor basis space model according to the order of the low-order sub-tensor corresponding to each multi-source heterogeneous data and the order corresponding to each basic attribute information included in the tensor basis space model to obtain a high-order tensor space model;
[0173] A unified representation of multi-source heterogeneous data is obtained based on a high-order tensor space model.
[0174] Optionally, the basic attribute information includes time information, space information, and user information;
[0175] The high-order tensor space model is obtained by the following formula:
[0176] X=T u ∪T semi ∪T s =(time,space,user,i1,…,i N ,V)
[0177] Among them, X represents the P+3 order tensor, time represents time information, space represents space information, user represents user information, i1,…,i N Indicates the order of the object that can be expanded, and V indicates the value.
[0178] Optionally, the data extraction module performs matrix singular value decomposition on the matrix using the following formula to obtain a corresponding decomposition result:
[0179]
[0180] Among them, × k ,k=1,2,…,N represents the modal product, is the order tensor model of the data to be extracted; A (1) ,A (2) ,…,A (N) are factor matrices of different dimensions, is the core tensor.
[0181] Optionally, if the data to be extracted is special data, which is at least one of data with incomplete attribute information, data with inconsistent attribute information, redundant data, and noise data in the multi-source heterogeneous data, the data extraction module is further configured to:
[0182] Determine the rank tensor model of special data from the higher-rank tensor space model;
[0183] Perform multi-level high-order singular value decomposition on the order tensor model of special data to obtain the data corresponding to the special data.
[0184] Optionally, the data extraction module performs multi-level high-order singular value decomposition on the order tensor model of special data using the following formula:
[0185] X=U1U2…U m W m +U s W s
[0186] Among them, X is the order tensor model, U1U2…U m are basis matrices of different orders, W s ,W m is the corresponding coefficient matrix.
[0187] Optionally, the device further includes an expansion module, specifically configured to:
[0188] Get the newly added data and determine the low-order sub-tensor corresponding to the newly added data;
[0189] If the high-order tensor space model does not include the order of the low-order sub-tensor corresponding to the newly added data, the excluded order is added to the high-order tensor space model as an expandable object order to obtain an expanded high-order tensor space model, and the newly added data is uniformly represented based on the expanded high-order tensor space model;
[0190] If the high-order tensor space model includes the order of the low-order sub-tensors corresponding to the newly added data, the same order is merged and the different orders are retained in the high-order tensor space model to obtain a unified representation of the newly added data.
[0191] Optionally, each order in the high-order tensor space model includes at least one dimension, and the extension module merges the same orders in the high-order tensor space model, specifically for:
[0192] In high-order tensor space models, the same order is merged using the finest-grained dimension.
[0193] The heterogeneous data access device of this embodiment can execute the heterogeneous data access method shown in the embodiment of this application. The implementation principle is similar and will not be repeated here.
[0194] An embodiment of the present application provides an electronic device, which includes: a processor; and a memory, wherein the memory is configured to store machine-readable instructions, which, when executed by the processor, enable the processor to execute a method for accessing heterogeneous data.
[0195] Compared with existing technologies, this embodiment of the present application proposes a unified high-order tensor space representation model based on the characteristics of attribute information of data with different structures. This model achieves a unified representation (i.e., a tensor model) for unstructured, semi-structured, and structured data, establishing a concise and efficient theoretical representation model for vehicle-road-cloud integrated big data analysis and processing. Furthermore, during subsequent data access, the unified representation model corresponding to the data to be accessed can be quickly determined based on the attribute information of different heterogeneous data, thereby achieving rapid data access.
[0196] The present application embodiment provides an electronic device, such as Figure 6 As shown, Figure 6 The electronic device 2000 shown includes a processor 2001 and a memory 2003. The processor 2001 and the memory 2003 are connected, for example, via a bus 2002. Optionally, the electronic device 2000 may further include a transceiver 2004. It should be noted that in actual applications, the number of transceivers 2004 is not limited to one, and the structure of the electronic device 2000 does not constitute a limitation on the embodiments of the present application.
[0197] Processor 2001 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 2001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0198] The bus 2002 may include a path for transmitting information between the above components. The bus 2002 may be a PCI bus or an EISA bus, etc. The bus 2002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0199] The memory 2003 may be a ROM or other type of static storage device that can store static information and instructions, a RAM or other type of dynamic storage device that can store information and instructions, or an EEPROM, a CD-ROM or other optical disk storage, an optical disc storage (including a compact disc, a laser disc, an optical disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0200] The memory 2003 is used to store the application code for executing the solution of the present application, and the execution is controlled by the processor 2001. The processor 2001 is used to execute the application code stored in the memory 2003 to implement Figure 5 The illustrated embodiment provides actions of a device for accessing heterogeneous data.
[0201] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0202] The above description is only a partial embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for accessing heterogeneous data, characterized in that: include: Obtaining a high-order tensor space model, the high-order tensor space model including a unified representation of multi-source heterogeneous data in vehicle-road-cloud integration, the high-order tensor space model being obtained by fusing low-order sub-tensors of each of the multi-source heterogeneous data into a tensor basis space model, the order of each of the low-order sub-tensors being determined based on attribute information of the corresponding multi-source heterogeneous data; the tensor basis space model including orders corresponding to at least three basic attribute information; the basic attribute information including time information, spatial information, and user information; Determining an order tensor model of the data to be extracted from the high-order tensor space model; Performing modular expansion on the order of the order tensor model to obtain a corresponding matrix; Performing matrix singular value decomposition on the matrix to obtain a corresponding decomposition result; Performing modular multiplication on the decomposition result to obtain the data to be extracted; If the data to be extracted is special data, and the special data is at least one of data with incomplete attribute information, data with inconsistent attribute information, redundant data, and noise data in each of the multi-source heterogeneous data, the method further includes: Determining an order tensor model of the special data from the high-order tensor space model; Performing multi-level high-order singular value decomposition on the order tensor model of the special data to obtain data corresponding to the special data; The multi-level high-order singular value decomposition is performed on the order tensor model of the special data using the following formula: X=U1U2…u m W m +U s W s Among them, X is the order tensor model, U1U2…U m are basis matrices of different orders, W s ,W m is the corresponding coefficient matrix.
2. The method according to claim 1, characterized in that The multi-source heterogeneous data includes at least one of real-time message data in vehicle-road-cloud integration, various structured report data, attribute data, unstructured text and image data, and various video and voice streaming data.
3. The method according to claim 1 or 2, characterized in that The unified representation of the multi-source heterogeneous data is obtained by the following method: Acquire each of the multi-source heterogeneous data and attribute information corresponding to each of the multi-source heterogeneous data; Determining, according to the attribute information corresponding to each of the multi-source heterogeneous data, a low-order sub-tensor corresponding to each of the multi-source heterogeneous data, wherein different orders of the low-order sub-tensors represent different attribute information; Obtaining the tensor basis space model; Arranging each of the multi-source heterogeneous data into the tensor basis space model according to the order of the low-order sub-tensor corresponding to each of the multi-source heterogeneous data and the order corresponding to each of the basic attribute information included in the tensor basis space model to obtain the high-order tensor space model; A unified representation of the multi-source heterogeneous data is obtained based on the high-order tensor space model.
4. The method according to claim 1, wherein The high-order tensor space model is obtained by the following formula: X=T u ∪T semi ∪T s =(time,space,user,i1,…,i N ,V) Among them, X represents the P+3 order tensor, time represents time information, space represents space information, user represents user information, i1,…,i N Indicates the order of the object that can be expanded, and V indicates the value.
5. The method according to claim 1, characterized in that Perform matrix singular value decomposition on the matrix using the following formula to obtain the corresponding decomposition result: Among them, × k ,k=1,2,…,N represents the modal product, is the order tensor model of the data to be extracted; A (1) ,A (2) ,…,A (N) are factor matrices of different dimensions, is the core tensor.
6. The method according to claim 1, characterized in that The method further comprises: Obtaining new data and determining the low-order sub-tensor corresponding to the new data; If the high-order tensor space model does not include the order of the low-order sub-tensor corresponding to the newly added data, the excluded order is added to the high-order tensor space model as an expandable object order to obtain an expanded high-order tensor space model, and the newly added data is uniformly represented based on the expanded high-order tensor space model; If the high-order tensor space model includes the order of the low-order sub-tensor corresponding to the newly added data, the same order is merged and the different orders are retained in the high-order tensor space model to obtain a unified representation of the newly added data.
7. The method according to claim 6, characterized in that Each order in the high-order tensor space model includes at least one dimension, and merging the same orders in the high-order tensor space model includes: In the high-order tensor space model, the same order is merged using the finest-grained dimension.
8. A device for accessing heterogeneous data, characterized in that: include: A model acquisition module for acquiring a high-order tensor space model, the high-order tensor space model including a unified representation of multi-source heterogeneous data in vehicle-road-cloud integration, the high-order tensor space model being obtained by fusing low-order sub-tensors of each of the multi-source heterogeneous data into a tensor basis space model, the order of each of the low-order sub-tensors being determined based on attribute information of the corresponding multi-source heterogeneous data; the tensor basis space model including orders corresponding to at least three basic attribute information; the basic attribute information including time information, spatial information, and user information; A model determination module, configured to determine an order tensor model of the data to be extracted from the high-order tensor space model; a data extraction module, configured to perform modular expansion on the order of the order tensor model to obtain a corresponding matrix, perform matrix singular value decomposition on the matrix to obtain a corresponding decomposition result, and perform modular multiplication on the decomposition result to obtain the data to be extracted; If the data to be extracted is special data, and the special data is at least one of data with incomplete attribute information, data with inconsistent attribute information, redundant data, and noise data in each of the multi-source heterogeneous data, the data extraction module is further used to: determine an order tensor model of the special data from the high-order tensor space model; perform multi-level high-order singular value decomposition on the order tensor model of the special data to obtain data corresponding to the special data; wherein the multi-level high-order singular value decomposition of the order tensor model of the special data is performed by the following formula: X=U1U2…U m W m +U s W s Among them, X is the order tensor model, U1U2…U m are basis matrices of different orders, W s ,W m is the corresponding coefficient matrix.
Citation Information
Patent Citations
Streaming data increment processing method and device based on tensor chain decomposition
CN111241076A
Cloud manufacturing resource servitization packaging and publishing method based on tensor theory
CN116232896A