LARGE-SCALE DATA STREAM PROCESSING DEVICE

The large-scale data flow processing device addresses the challenge of managing and processing large data flows by employing a modular platform with semantic reasoning and decision rules, enabling efficient real-time data management and adaptive deployment.

FR3061577B1Active Publication Date: 2025-09-12BULL SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2016063537
Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2016-12-30
Publication Date
2025-09-12
Estimated Expiration
2036-12-30

AI Technical Summary

Technical Problem

Current technologies are inadequate for managing and processing large volumes of data flows, particularly 'hot' data that changes in real-time, requiring new approaches to handle the intersection and interoperability of data streams.

Method used

A large-scale data flow processing device comprising a knowledge base, front-end communication device, and decision-making device, with modular platform architecture, capable of capturing, processing, and storing data streams, and applying semantic reasoning and decision rules to produce actionable insights.

Benefits of technology

Enables efficient management and processing of large data flows, allowing for real-time decision-making and adaptive deployment on cloud-ready virtual machines, optimizing network traffic, and supporting dynamic data visualization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000013_0000
    Figure 00000013_0000
Patent Text Reader

Abstract

The present invention relates to a device for processing large-scale data flows (big data), comprising a knowledge base, a hardware and software assembly constituting a front-end communication device, making it possible to capture flows from the external environment and, if necessary, to return data to this environment, the front-end delivering the flows to a platform to undergo various processing operations there, traces are collected and stored in a storage and memorization architecture during the execution of the processing operations, the platform producing data which feeds a decision-making device comprising a hardware and software assembly defining decision rules making it possible either to trigger actions or to initiate feedback to the front-end communication device or to the knowledge base of said processing device. Abstract figure: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Device for processing large-scale data streams TECHNICAL FIELD OF THE INVENTION

[0001] The present invention relates to the field of data processing and more specifically to a device for processing large volumes of data. TECHNOLOGICAL BACKGROUND OF THE INVENTION

[0002] Data is everywhere the raw material of our world, which is now definitively digital. The information deluge creates new opportunities, for example, for business. This mass data redefines the way scientific knowledge is created and also offers companies new levers for growth.

[0003] The intersection of data flows that now irrigate a large number of sectors, in particular the economic domains, must allow access to systemic information inaccessible in each data flow taken in isolation. The Semantic Web offers a framework for implementing this intersection.

[0004] Semantic web standards are beginning to be well disseminated and stabilized (RDF, OWL for the representation of data and metadata, as well as protocols for exchanges, essentially HTTP). They were created to facilitate interoperability and the exchange of data. The Web has effectively become the preferred source of data and the place of the most dynamic exchanges. The provision and promotion of open or semi-open public data, their combination with industrial data, and the tools to exploit them are gradually giving rise to more than one important lever to boost, for example, the economic sector.

[0005] The work of collecting, integrating, analyzing, using, and visualizing data is becoming systematic among many stakeholders. Today, it largely concerns "cold" data or data that changes little over time. But a new interest is emerging for "hot" data, i.e. data close to real time, which poses new problems and requires new approaches.

[0006] Some components or tools for processing data streams, open source or commercial, already exist. This is the case, for example, of triple stores and SPARQL query execution engines (query language).

[0007] However, with the current volumes of data flows, their number and their variety, current techniques and tools are no longer able to meet the requirements of users. GENERAL DESCRIPTION OF THE INVENTION

[0008] The present invention aims to overcome certain drawbacks of the prior art concerning the management and processing of large data flows.

[0009] This aim is achieved by a large-scale data flow processing device (big data), comprising a knowledge base, a hardware and software assembly constituting a front-end communication device, making it possible to capture flows from the external environment and, if necessary, to return data to this environment, the front-end delivering the flows to a platform to undergo various processing operations, traces being collected and stored in a storage and memorization architecture during the execution of the processing operations, the platform produces data which feeds a decision-making device comprising a hardware and software assembly defining decision rules making it possible either to trigger actions or to initiate feedback to the front-end communication device or to the knowledge base.

[0010] According to another particularity, the different treatments include the semantics of the flow, or the creation of a summary, or the crossing of at least two flows or the interconnection of several flows between them.

[0011] According to another particularity, the platform comprises a plurality of “smart'op” which are flow processing processes, these processes being for example scripted, via a scripting means, to produce data feeding the decision-making device.

[0012] According to another feature, the storage of activity traces in the storage architecture allows the platform, by accessing these activity traces, to carry out repetitions of possible games or deferred processing.

[0013] According to another particularity, the semantics of the flows conforms to an ontology of description of the flows.

[0014] According to another feature, the interconnection of the flows can include semantic reasoning.

[0015] According to another particularity, all the operations to which the flows are subjected contribute to the production of data.

[0016] According to another particularity, the data can be alerts or singularities resulting from the reasoning carried out on the flows.

[0017] According to another feature, the knowledge base contains information representing knowledge of the external environment and includes the description of the sensors and the infrastructure which supports them.

[0018] According to another particularity, this information constituting knowledge can be modified by feedback triggered by decision rules (housed in the “decision maker”). PRESENTATION OF ILLUSTRATIVE FIGURES

[0019] Other features and advantages of the present invention will appear more clearly on reading the following description, given with reference to the appended drawings, in which:

[0020] - [Fig.l] represents the diagram of a data flow processing system comprising at least one device and a stream processing platform, according to one embodiment;

[0021] DESCRIPTION OF THE PREFERRED EMBODIMENTS OF THE INVENTION

[0022] The present invention relates to a device for processing large data streams.

[0023] In certain embodiments, the large-scale data flow processing device (big data) comprises a knowledge base (1, [Fig.l]), a hardware and software assembly constituting a front-end communication device (2), making it possible to capture the flows from the external environment (3) and, if necessary, to restore data to this environment, the front-end delivering the flows to a platform (4) to undergo various processing operations there, activity traces are collected and stored in a storage and memorization architecture (5) during the execution of the processing operations, the platform (4) producing data which feeds a decision-making device (6) comprising a hardware and software assembly defining decision rules making it possible either to trigger actions (7) or to initiate feedback (8) to the front-end communication device (2) or to the knowledge base (1) of said device.

[0024] The data flow processing device and the platform (4) allow users to efficiently manage their own data flows and those produced in their field of activity by other actors (partners, suppliers, customers, organizations, agencies, etc.). These data also include open data (Open Data), linked (Linked Open Data) or not, which are produced by international communities or private organizations.

[0025] The language or format or model used for the description of the data in the platform (4) is preferably RDF (Resource Description Framework). The RDF (graph) model allows the formal description of Web resources and their metadata, so as to allow the automatic processing of such descriptions. A document structured in RDF is a set of triplets. An RDF triplet is an association (subject, predicate, object): • the “subject” represents the resource to be described; • the “predicate” represents a type of property applicable to this resource; • The “object” represents a piece of data or other resource, it is the value of the property.

[0026] The subject, and the object in the case where it is a resource, can be identified by a URI (uniform resource identifier) ​​or be anonymous nodes. The predicate is necessarily identified by a URI.

[0027] RDF documents can be written in various syntaxes, including XML. But RDF itself is not an XML dialect. Other syntaxes can be used to express triples. RDF is simply a data structure made up of nodes and organized as a graph. Although RDF / XML—its XML version proposed by the W3C (World Wide Web Consortium)—is only a syntax (or serialization) of the model, it is often referred to as RDF. A misnomer refers to both the triple graph and the XML presentation associated with it.

[0028] An RDF document thus formed corresponds to a labeled directed multi-graph. Each triplet then corresponds to a directed arc whose label is the predicate, the source node the subject and the target node the object.

[0029] The description of documents and / or data in RDF is generally based on one or a set of ontology(ies). An ontology is a structured set of terms and concepts representing the meaning of a field of information, whether through the metadata of a namespace, or the elements of a knowledge domain. The ontology itself constitutes a data model representative of a set of concepts in a domain, as well as the relationships between these concepts. It is used to reason about the objects of the domain concerned.

[0030] The concepts are organized in a graph and are linked to each other by taxonomic relationships (hierarchy of concepts) on the one hand, and semantic relationships on the other.

[0031] This definition makes it possible to write languages ​​intended to implement ontologies. To construct an ontology, we have at least three of these notions: • determination of passive or active agents; • their functional and contextual conditions; • their possible transformations towards limited objectives.

[0032] To model an ontology, we will use these tools to: • refine adjacent vocabularies and concepts; • break down into categories and other topics; • predict in order to know the adjacent transformations and orient towards the internal objectives; • relativize in order to encompass concepts; • similarize in order to reduce to completely distinct bases; • instantiate in order to reproduce the whole of a “branch” towards another ontology.

[0033] Ontologies are used in artificial intelligence, the Semantic Web, software engineering, biomedical computing, and information architecture as a form of knowledge representation about a world or a certain part of that world. Ontologies generally describe: • individuals, constituting the basic objects; • classes: constituting sets, collections, or types of objects; • attributes: consisting of properties, functionalities, characteristics or parameters that objects can possess and share; • relationships, constituting the links that objects can have between them; • events representing changes to attributes or relationships; • metaclass (semantic web), constituting collections of classes which share certain characteristics

[0034] In some embodiments, the different treatments include semantization of the flow, or creation of a summary, or crossing of at least two flows or interconnection of several flows between them.

[0035] In certain embodiments, the platform (4) comprises a plurality of “smart'op” which are flow processing processes, these processes being scripted, via a scripting means, for example and in a non-limiting manner a DSL (domain-dedicated language), to produce data feeding the decision-making device (6).

[0036] Scripting means programming using scripts. A script is defined as a program in interpreted language.

[0037] In certain embodiments, the storage of activity traces in the storage architecture allows the platform (4) by accessing these traces to perform repetitions of possible games or deferred time processing.

[0038] By repetition of games or replay, we mean the operation which consists of “redoing” a previously carried out processing and of which the activity traces have been kept.

[0039] In some embodiments, the semantics of the flows conform to a flow description ontology.

[0040] In some embodiments, the interconnection of the streams may include semantic reasoning.

[0041] In certain embodiments, all the operations to which the flows are subjected (crossing the platform (4)), contribute to the production of data (input flow of the “decision-maker” (6), see [Fig. 1]).

[0042] In some embodiments, the data may be alerts or singularities resulting from reasoning performed on the streams.

[0043] In some embodiments, the knowledge base contains information representing knowledge of the external environment and includes the description of the sensors and the infrastructure that supports them.

[0044] In some embodiments, this information constituting the knowledge can be modified by feedback (8) triggered by decision rules (housed in the “decision maker” (6)). For example and in a non-limiting manner, the “decision maker” (6), discovering that a sensor is defective, updates the knowledge base. The “decision maker” (6) can also transmit data to the external environment (3) by communicating with the “front end” (2), as shown in [Fig.l].

[0045] In certain embodiments, the platform (4) has a modular architecture, to support the addition or replacement of (new) components without recompilation, and to allow deployment on a fleet of standardized virtual machines ("cloud-ready"). Said platform (4) is modularized by means of abstraction layers making it possible to encapsulate all access to the services of the operating system (file system), storage (database, RDF triple store, etc.) or exchange (messaging bus).

[0046] This modular platform architecture can rely on market standards to integrate into user systems, in particular: • The Java platform and its ecosystem. Java-based platforms benefit from many advantages. The execution engine, running on several operating systems, Java developments inherit platform independence (4) and code portability, which then runs identically on each of them. It is an environment where function libraries and frameworks are available in open source, which facilitates development. Beyond the programming language, Java is above all a deployment platform (the Java virtual machine, the library and application servers of the enterprise version (Java EE) which supports several types of programming: imperative (Java), dynamic / scripted (Groovy, JRuby...) and functional (Scala, Closure...) • W3C (World Wide Web Consortium) standards, whether at the level of data modeling (RDF, XML), query formulation (querying using SPARQL language), infrastructure (HTTP protocol, REST model) or presentation (HTML 5, CSS)

[0047] In certain embodiments, a replacement of the implementations of said abstractions makes it possible to move from simple implementations adapted to development (local file system, triple store in memory, event observers) to others focused on large-scale deployment (virtual machine park, cluster or server farm): distributed file systems of the “Grid” type (MongoDB, Infinispan, etc.), distributed or clustered triple store, JMS or AMQP messaging bus, etc.,

[0048] In certain embodiments, the architecture of the platform (4) makes it possible to adapt the deployment of the system to the target volume but also to make it evolve in the event of evolution, by simply adding virtual machines if the selected components allow it or, by replacing the components limiting the scalability.

[0049] Another aspect of system scalability concerns the management of network flows and the optimization of traffic to or from the system. The use of the REST (REpresentational State Transfer) model for the provision of services to applications and users makes it possible to optimize the use of the network infrastructure, both on the server side and on the client side, for example through cache management directives (servers, networks (proxies) and client) but also to avoid unnecessary requests such as content revalidation.

[0050] Similarly, fine-grained management of content negotiation makes it possible to optimize exchanges by minimizing or even eliminating the number of redirections necessary to guide the user to the resource providing content adapted to their constraints / capacities (MIME types). The architecture can support the interfaces necessary for this optimization of the exchanged flows.

[0051] Beyond the technical aspects, this task makes it possible to establish the rules for routing and distributing events concerning the flows in order to automate and optimize the distribution of processing according to the topology (architecture) of the platform (4) and the processing capacity of each smarfop instance composing it. Said architecture also allows a dynamic and self-adaptive distribution of processing, capable of taking into account events such as the disappearance of a node (recovery) and its return to service, or even, optionally, the introduction of a new node (hot scalability).

[0052] In some embodiments, the SPARQL query language is extended to perform summaries of data streams or implement forget functions to historicize data streams with variable granularity over time and ensuring limited storage space.

[0053] The summaries also make it possible to specify and implement semantic filtering operators (because they are applied to semantic data) which use as input the user's context which is reflected by their profile, their data access rights, their geolocation, their preferences, their terminal as well as their environment (weather, season, etc.)

[0054] The dynamic data recovered from the various sensors and other flows are semantized. This semantization results in a conversion of this data into RDF triplets enhanced with a temporal dimension which characterizes the almost continuous arrival of these flows.

[0055] Thus, the extension of the SPARQL language, by integrating notions such as the time window, makes it possible to make queries, filter or reason on these semantic flows.

[0056] When a system requires fast and intelligent processing of large amounts of data that arrive continuously, it becomes very expensive and sometimes redundant or impossible to store all the streams before processing them. It is therefore necessary to process the semanticized data on the fly and to store only the relevant data by making summaries (for example, extracting a representative sample of the stream using statistical approaches).

[0057] In some embodiments, the SPARQL language is extended by introducing the notion of an adaptable time window (defined portion of a stream) and operators specific to stream processing. Queries processing the data must adapt to the arrival rate of dynamic data and be continuously evaluated in order to take into account the evolving nature of the stream. The semantics of SPARQL queries thus allow processing based on the time or order of arrival of the data.

[0058] In some embodiments, the extended SPARQL language allows data that may be either static or dynamic and temporal to be mixed by interconnecting them, regardless of their number, source or quality.

[0059] In the context of the Semantic Web, the facts stored in knowledge bases are naturally not ordered. However, the streaming environment, with its notion of a time window for data processing, requires this support. The mechanisms for managing and representing incoming facts and for reasoning are therefore adapted to order said facts. Considerations of data velocity and volume imply an optimization of the aforementioned operations. Thus, it is essential to receive raw data, to semanticize them and to exploit them within an inference mechanism in a limited time, even if the data arrive at a very high speed and in very large volumes. Taking these two factors into account with the time constraint of the flow analysis window is ensured by a dynamic system for their intermediate storage.Thus, in certain situations, the streams can be saved only in main memory while in other situations, it will certainly be necessary to persist them, even temporarily, in secondary memory.

[0060] For the exploitation of key / value type databases which are suitable for storage in main memory and offer a high level of performance, it is appropriate to adapt the RDF triples model to the key / value approach taken in its generality.

[0061] For the conversion of data into RDF triplets, a first approach called "direct mapping" makes it possible to automatically generate RDF from table names for classes and column names for properties. This approach makes it possible to quickly obtain RDF from a relational database, but without using a vocabulary.

[0062] The other approach provides the R2RML mapping language to associate vocabulary terms with the database schema. In the case of XML, a generic XSLT transformation can be performed to produce RDF from a wide range of XML documents.

[0063] Data visualization (dataviz) is one of the keys to their use. It allows us to understand, analyze, track, and detect key elements of the data. Above all, it allows the user to interact with the data, to perceive it physically. Much scientific work on data visualization has been and is being conducted in research laboratories. Many open source components allow an expert or semi-expert audience to grasp it and communicate their point of view on the data.

[0064] For example, data journalism (a movement aimed at renewing journalism by exploiting statistical data and making it available to the public) is booming due to two major factors. First, the public availability of data (Open Data) allows rapid access to trusted data. Some of this data is accessible in the form of streams, but this movement is tending to grow thanks to the arrival of the "Internet of Things" and smart cities. This will allow access to data in streams of data from sensors or other communicating objects. Second, the ability to implement attractive visualizations has made it possible to make the discourse and argumentation clearer and more meaningful. The visual representation of static or relatively static data has reached a first degree of maturity.The use of the extended SPARQL language allows the modular platform (4) to also handle the visualization of dynamic data and semantic flows.

[0065] The present application describes various technical features and advantages with reference to the figures and / or to various embodiments. Those skilled in the art will understand that the technical features of a given embodiment may in fact be combined with features of another embodiment unless the opposite is explicitly mentioned or it is obvious that these features are incompatible or the combination does not provide a solution to at least one of the technical problems mentioned in the present application. In addition, the technical features described in an embodiment given can be isolated from other features of that mode unless the opposite is explicitly mentioned.

[0066] It should be obvious to those skilled in the art that the present invention allows embodiments in many other specific forms without departing from the scope of the invention as claimed. Therefore, the present embodiments should be considered illustrative, but may be modified within the scope defined by the protection sought, and the invention should not be limited to the details given above.

Claims

Claims

1. Device for processing large-scale data flows, comprising a knowledge base (1) which includes information representing knowledge of the external environment (3) and the description of data flow sensors and the infrastructure supporting them, a hardware and software assembly constituting a front-end communication device (2), making it possible to capture the data flows from the external environment (3) and, if necessary, to return data to this environment (3), the front-end delivering the data flows from said external environment to a platform (4), the latter interrogating said knowledge base (1) and carrying out various processing operations on the data, traces being collected and stored in a storage and memorization architecture (5) during the execution of the processing operations by the platform (4) on the data from the external environment (3),the processing device being characterized in that it comprises a decision-making device (6) monitoring the data produced by the platform (4) and comprising a hardware and software assembly defining decision rules allowing either the transmission of the data to the external environment by communicating with the front-end device (2) or the updating of the knowledge base (1) of said processing device.,

2. Device according to claim 1, characterized in that the different treatments comprise at least one of the following treatments: the semantics of the flow, the creation of a summary, the crossing of at least two flows, the interconnection of several flows between them.

3. Device according to claim 1 or 2 characterized in that the platform (4) comprises a plurality of “smart'op” which are flow processing processes, these processes being scripted, via a scripting means, to produce data feeding the decision-making device (6).

4. Device according to claim 1 characterized in that the platform (4) accesses traces stored in the storage architecture and performs repetitions of possible games or deferred time processing based on said traces.

5. Device according to claim 2 characterized in that the semantics of the flows conforms to an ontology of description of the flows.

6. Device according to claim 2 characterized in that the interconnection of the flows includes semantic reasoning.

7. Device according to one of claims 1 to 6, characterized in that the data are alerts or singularities resulting from the reasoning carried out on the flows.

8. Device according to claim 1, characterized in that the architecture of the platform (4) is modularized by means of abstraction layers which encapsulate all access to the services of the operating, storage or exchange system to support the addition or replacement of components of said platform (4) without recompilation, and to carry out the deployment on a fleet of standardized virtual machines.

9. Device according to claim 8, characterized in that the modular architecture of the platform (4) allows the transition from simple implementations adapted to development to others focused on large-scale deployment by replacing the implementations of said abstraction layers of said modular architecture.