DATA STRUCTURE AND CORRESPONDING IMPLEMENTATION PROCEDURE

DE602020072845T2Active Publication Date: 2026-06-03BULL SA

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
BULL SA
Filing Date
2020-12-14
Publication Date
2026-06-03
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

DATA STRUCTURE AND ASSOCIATED IMPLEMENTATION PROCESS Technical field

[0001] The invention relates to the field of data processing equipment and methods specifically adapted to particular functions. In particular, it relates to a method for converting a first data file, in a first format, into a second data file in a second generic format different from the first. The invention also relates to a storage medium and a data structure associated with the method. Previous technique

[0002] We are familiar with computer platforms for data aggregation.

[0003] Such computer platforms generally fulfill three missions. The first mission, known as investigation, allows for the reconstruction of the past actions (movements, contacts, etc.) of suspects linked to one or more events under investigation. The second mission, known as monitoring, provides continuous information on the current activities of groups of people suspected of engaging in illegal activities. The third mission, known as surveillance, relates to the detection of early warning signs of events or indicators of illegal activities. The communication and location data associated with these platforms constitute an increasingly important source of information.

[0004] In order to function, such computer platforms require importing data from various sources.

[0005] Due to the many different source formats that exist, the processing units of these computer platforms do not have the capacity to process all types of data produced.

[0006] Thus, it is known that some data is simply ignored because it is unknown to these computer platforms.

[0007] The document US2008243891A1 discloses the use of correspondence relationships between data formats for conversion in a network. Summary of the invention

[0008] The invention aims to overcome this drawback.

[0009] The invention relates in particular to a method for converting a first data file, called the conversion source file, into a second data file, called the conversion destination file. In the invention, the conversion source file has a predetermined proprietary format specific to any one of a plurality of data sources. Furthermore, the conversion destination file has a predetermined generic format independent of any one of the plurality of data sources. In addition, the predetermined generic format is associated with a predetermined schema that describes the semantic structure of the predetermined generic format. The method comprises the following steps: a step of obtaining correspondence relations between the format of the original conversion file and the format of the destination conversion file, a step of forming the destination conversion file from the predetermined schema and the data from the original conversion file that are involved in the correspondence relations, a step of detecting if there remains data from the original conversion file that is not involved in the correspondence relations, called unmatched data, and in this case, a step of extending the format of the destination conversion file to include all or part of the unmatched data.

[0010] In a first implementation, the data in the original conversion file relates to at least one discursive electromagnetic signal that indicates the implementation of at least one communication service in at least one telecommunication network via at least one communication medium, the data describing the type of service and / or communication medium, the start time and possibly the end time of the communication service and one or more nodes of the telecommunication network involved in the implementation of the discursive communication service.

[0011] In a second implementation, the data includes information from the interception of telecommunications, known as "COMINT" domain data.

[0012] In a third implementation, the predetermined schema includes a plurality of fields capable of describing the largest number of data associated with the plurality of data sources.

[0013] In a fourth implementation, the step of obtaining correspondence relationships includes a step of analyzing at least one intermediate file that defines predefined correspondence relationships between at least one keyword associated with the content of the plurality of data sources and all or part of the fields of the predetermined schema format.

[0014] In an example of the fourth implementation, the intermediate file analysis step includes, a step of calculating a degree of correspondence between the keyword associated with the content of the plurality of data sources and the fields of the predetermined schema format, and a step of determining the field of the predetermined schema format having a degree of correspondence greater than a predefined degree of correspondence threshold, as the target field of the predetermined schema format corresponding to the keyword.

[0015] In a fifth implementation, the predetermined schema is an XSD schema and the conversion destination file formation step includes the formation of a file in XML format.

[0016] In a sixth implementation, the predetermined schema is a JSON schema and the conversion destination file formation step includes the formation of a file in JSON format.

[0017] In a seventh implementation, when it depends on the second implementation, the training step includes a step of identifying, from the original conversion file, at least one unique path taken by a communication service that has passed through the node(s) of the telecommunication network involved in the implementation of the discursive communication service.

[0018] In this implementation, the predetermined scheme is arranged so as to separate, on the one hand, the node(s) of the telecommunications network involved in the implementation of the discursive communication service and, on the other hand, the unique path(s) taken by a communication service that has passed through the node(s) of the telecommunications network involved in the implementation of the discursive communication service.

[0019] In an eighth implementation, the step of extending the format of the conversion destination file includes extending a predetermined field of the predetermined schema.

[0020] The invention also relates to a recording medium recording an instruction program that can be executed by a processor to carry out all the steps of the process according to any one of the preceding implementations.

[0021] The invention also relates to a data structure with a predetermined generic format and independent of any data source for storing data, the data structure being obtained by a conversion process according to any of the previous implementations from a data file and a predetermined schema which defines the structure of the predetermined generic format.

[0022] In an initial implementation, the predetermined schema is an XSD schema and the data structure is an XML file.

[0023] In a second implementation, the predetermined schema is a JSON schema and the data structure is a JSON file.

[0024] In a third implementation, the structure further comprises a first part of the data structure and a second part of the data structure which are separated from each other, the first part of the data structure comprising a list of the node(s) of the telecommunications network involved in the implementation of the discursive communication service, the second part of the data structure comprising a list of the unique path(s) taken by a communication service which has transited through the node(s) of the telecommunications network involved in the implementation of the discursive communication service. A brief description of the designs

[0025] Other features and advantages of the invention will be better understood from the description that follows and with reference to the attached drawings, given for illustrative purposes only and not for limitation.

[0026] [ Fig. 1 ] There figure 1represents a method for implementing the invention.

[0027] The single figure does not necessarily respect scales, particularly in thickness, and this is for illustrative purposes. Description of the methods of realization

[0028] The invention relates to a generic and extensible data structure for transferring data between communication interception devices and a chain for processing intercepted information / communications.

[0029] The purpose of this invention is to cover all communication interception devices, whether they are dedicated to intercepting telephone communications or internet streams.

[0030] More specifically, the invention relates to a method of converting a first data file, called the conversion source file, into a second data file, called the conversion destination file.

[0031] In a non-limiting example, the data according to the invention are intelligence data. However, the invention can be applied to other types of data that need to be converted from a first format to a generic format.

[0032] In a particular implementation, the data relates to at least one discursive electromagnetic signal that indicates the implementation of at least one communication medium in at least one known type telecommunication network via at least one communication service.

[0033] A discursive electromagnetic signal is defined as an electromagnetic signal whose purpose is to transmit information between at least two devices.

[0034] For example, an electromagnetic signal associated with communication established according to a given telecommunication protocol (e.g., phone calls, WhatsApp, Signal or Skype; SMS, MMS messages) is of the discursive type, because such an electromagnetic signal allows information to be transmitted between at least two devices.

[0035] In contrast, for example, a radar signal is a non-discursive signal, as its purpose is not to transmit information between at least two devices. Indeed, as is well known, radar is a remote sensing tool that detects objects at a distance. To do this, the radar emits waves and then compares the waves reflected off these objects to obtain information about the objects' distance, speed, and so on.

[0036] Communication medium refers to one of the methods of exchanging or disseminating information that are generally implemented by a telecommunications network.

[0037] A communication service is defined as a means that allows users of a telecommunications network to use one or more communication media.

[0038] Thus, among other things, a communication medium can be chosen from an audio call, a video call, a text message, a multimedia message, an email, or a web session. A communication service may allow the use of one or more communication media. For example, the Skype service allows the use of at least audio, video, and text messaging.

[0039] In one particular implementation, the data includes information from the interception of telecommunications, known as data from the "COMINT" domain ("Communications Intelligence" in English; intelligence from the interception of telecommunications, in French).

[0040] In a first example, COMINT data describes one or more nodes of the telecommunications network involved in the implementation of the discursive communication service.

[0041] In practice, in this example, COMINT data corresponds to identifiers of equipment and / or users of that equipment.

[0042] For example, equipment can be a mobile phone, a landline phone, a computer, a tablet, a computer server, or network equipment such as a switch, router, base station, or radio network controller.

[0043] Thus, without limitation, an identifier can be an IP address (“Internet Protocol”), a MAC address (“Media Access Control”), an MSISDN telephone number (“Mobile Station ISDN Number”), an IMEI terminal identifier (“International Mobile Equipment Identity”), an IMSI subscription identifier (“International Mobile Subscriber Identity”), a BSIC base station identifier (“Base Station Identity Code”).

[0044] In a second example, the COMINT data describes the type of communication medium and / or service used. As mentioned above, the type of communication medium can be chosen from an audio call, a video call, a text message, a multimedia message, an email, or a web session, while the type of communication service can be chosen from WhatsApp, Signal, or Skype.

[0045] In a third example, COMINT data describes the start time and possibly the end time of the communication service.

[0046] Thus, typically, a video call, an audio call, and a web session have a start time and an end time. However, a text message, a multimedia message, and an email only have a start time.

[0047] In a fourth example, COMINT data describes the successive positions of a mobile communication device, either in the form of a direct location, or in the form of an indirect, and more imprecise, location through the locations of the relay equipment of a telecommunications network to which the mobile device has connected.

[0048] In the invention, the original conversion file has a predetermined proprietary format specific to any one of a plurality of data sources.

[0049] Typically, a data source is a sensor designed to intercept communications within a telecommunications network. Such sensors can belong to various third-party entities, such as telecommunications operators or a state's internal or external security forces. Furthermore, in practice, each data source generates data files in one or more formats specific to it.

[0050] Also in the invention, the conversion destination file has a predetermined generic format independent of any one of the plurality of data sources.

[0051] One objective of using a generic format is to ensure the independence of computer processing methods from the formats of the original files being converted. Another objective is to aggregate data from different sources. A further objective is to facilitate the exchange and sharing of data received from various sources between different organizations.

[0052] In the invention, the predetermined generic format is associated with a predetermined schema that describes the semantic structure of the predetermined generic format. In other words, the predetermined schema is configured to specify the content of the predetermined generic format.

[0053] In one particular implementation, the predetermined schema includes a plurality of fields capable of describing the largest number of data associated with the plurality of data sources.

[0054] In practice, this mainly involves covering data sources that operate in the fields of telephone communications and internet communications.

[0055] Thus, the predetermined schema allows for the description of most of the data produced by known data sources. This includes, but is not limited to, the identifier of an associated intercept sensor, the production date of the original conversion file, and the useful content of the data intercepted by the intercept sensor. For example, the useful content might include information such as: the characterization of the discursive electromagnetic signal (e.g. start and end date of the signal, identifiers of the devices at its origin and reception, identifiers of the communication medium(s), identifiers of the communication service); the informational content carried by the signal (e.g. text, audio, video); the characterization of the devices transporting the discursive electromagnetic signal and the signal paths through these devices; and the location of the transmitting or receiving devices.

[0056] In a first example, the predetermined schema is an XSD schema (“XML Schema Definition” in English; XML definition schema, in French).

[0057] In a second example, the predetermined schema is a JSON schema ("JavaScript Object Notation"; a textual data exchange language based on JavaScript).

[0058] In the example of the figure 1, process 100 first includes a step 110 of obtaining correspondence relations between the format of the original conversion file and the format of the destination conversion file.

[0059] In a particular implementation, the step 110 of obtaining correspondence relations includes an analysis step 111 of at least one intermediate file which defines predefined correspondence relations between at least one keyword associated with the content of the plurality of data sources and all or part of the fields of the predetermined schema format.

[0060] In one example, analysis step 111 of the intermediate file includes, a calculation step 1111 of a degree of correspondence between the keyword associated with the content of the plurality of data sources and the fields of the predetermined schema format, and a determination step 1112 of the field of the predetermined schema format which has a degree of correspondence greater than a predefined degree of correspondence threshold, as the target field of the predetermined schema format which matches the keyword.

[0061] Then, process 100 includes a step 120 of training the conversion destination file from the predetermined schema and data from the conversion source file which are involved in the matching relationships.

[0062] In a particular implementation, training step 120 includes an identification step 121, from the original conversion file, of at least one unique path taken by a communication service that transited through the telecommunication network node(s) involved in the implementation of the discursive communication service. Furthermore, in this implementation, the predetermined schema is arranged so as to separate, on the one hand, the telecommunication network node(s) involved in the implementation of the discursive communication service and, on the other hand, the unique path(s) taken by a communication service that transited through the telecommunication network node(s) involved in the implementation of the discursive communication service.

[0063] Thanks to this particular arrangement, it is easy to represent any transport graph of a communication service between a sender and at least one receiver involved in the communication service. For example, a graph can represent the different transport paths of a communication service sent to multiple recipients. In another example, a graph can represent the different transport paths of data packets for the same communication service.

[0064] In a first example, when the predetermined schema is an XSD schema, the 120 formatting step of the conversion destination file includes the formation of a file in XML format ("Extensible Markup Language" in English; extensible markup language, in French).

[0065] In a second example, when the predetermined schema is a JSON schema, the conversion destination file training step 120 includes training a file in JSON format.

[0066] Next, process 100 includes a detection step 130 to see if there is any data left in the original conversion file that is not involved in the matching relationships.

[0067] In other words, the goal here is to determine whether the original conversion file contains data that does not correspond to at least one of the fields in the predefined schema. This situation can occur when using a new data source or when modifying the data source that originally generated the original conversion file. For example, such a modification might be due to the data source adding data because of an update to the associated intercept sensor or the use of new intercept mechanisms.

[0068] In the event that there remains data from the original conversion file which is not involved in the matching relationships, called unmatched data, process 100 includes a step 140 of extending the format of the destination conversion file to include all or part of the unmatched data.

[0069] By extension, within the framework of the invention, we mean the possibility of defining a new type of data from a specific type of data from a source, unknown to the predetermined generic format.

[0070] This mechanism allows for the inclusion of specific data and new areas of communication interception. Furthermore, within the framework of data aggregation platforms, this mechanism avoids the need to develop data processing logic specific to each data source.

[0071] In one particular implementation, the extension step 140 of the conversion destination file format includes the extension of a predetermined field of the predetermined schema.

[0072] Thus, in this particular implementation, the extension consists of enriching a data type by adding content and / or attributes.

[0073] In computer science, schema extension is a way to obtain a derived type from a base type. The general idea of ​​extension is to add content and / or attributes from another type, either predefined or already defined in the schema. In other words, extension is similar to type derivation in object-oriented programming languages ​​like Java or C++.

[0074] For example, when the predetermined schema is of type XML, a type is introduced by the element "xsd:extension," whose attribute 'base' specifies the name of the base type. This can be a predefined type within the schema. The content of the "xsd:extension" element specifies the content and attributes to be added to the base type. The "xsd:extension" element is a child of an "xsd:simpleContent" or "xsd:complexContent" element, which is itself a child of the "xsd:complexType" element.

[0075] The following example corresponds to an extract from an XML file conforming to a predetermined schema according to the invention and which contains information related to an SMS communication. In this example, it is the tag ' <extensions>' which characterizes the extension according to the invention: <entry date="2020-10-10T12:00:42.395+02:00"> <commonkernel> <signalheader support="SMS" Service="Android Phone" volume="5000" duration="0" encrypt="CLAIR" / > <sender id="0666666666" type="phone" / > <receiver id="9292929492" type="tablet" / > <communicationnodes> <node id="2945388" name="NANTES2817727" / > <node id="1822213" name="PAU12" / > < / communication Nodes> <communicationpaths> <nodeevent id="2945388" date="2020-10-10T12:29:42.395+02:00" / > <nodeevent id="1822213" date="2020-10-10T12:12:00.200+02:00" / > < / communicationpaths> <content> There, refreshed by food and quiet, after the fear had passed, they attacked the wealthy villages with the arrival of cavalry cohorts, which happened to be approaching, and without attempting to resist the extended plain, they retreated and, yielding all the strength of their youth left, they took refuge in their settlements.< / content> < / communicationnodes> < / commonkernel> <extensions> <topic name="MessageClass"> <property name="className" value="flashMessage" / > <property name="classValue" value="0" / > < / topic> <topic name="Originatorlnfo"> <property name="OriginatorIMSI" value="0228458288" / > <property name="Originator MSISDN" value="033+02828828" / > < / topic> <topic name="Recipientinfo"> <property name="RecipientIMSI" value="01339929280" / > <property name="Recipient MSISDN" value="033+0282945228" / > < / topic> < / extensions> < / entry>

[0076] The invention also relates to a recording medium which records a program of instructions which can be executed by a processor to carry out all the steps of process 100 as described above.

[0077] This includes, but is not limited to, a hard drive, a CD-ROM, a DVD, a floppy disk, a cassette, or a USB flash drive.

[0078] The invention also relates to a data structure with a predetermined generic format. This generic format is independent of any data source, for storing data.

[0079] In the invention, the data structure is obtained by a method 100 as described above, from a data file and a predetermined schema which defines the structure of the predetermined generic format.

[0080] In a particular implementation, the data structure comprises a first part of the data structure and a second part of the data structure which are separated from each other.

[0081] In particular, the first part of the data structure includes a list of the node(s) of the telecommunications network involved in the implementation of the discursive communication service.

[0082] Furthermore, the second part of the data structure includes a list of the unique path(s) taken by a communication service that has passed through the node(s) of the telecommunication network involved in the implementation of the discursive communication service.

[0083] Thanks to this particular arrangement, it is easy to represent any transport graph of a communication service between a sender and at least one receiver involved in the communication service. For example, a graph can represent the different transport paths of a communication service sent to multiple recipients. In another example, a graph can represent the different transport paths of data packets for the same communication service.

[0084] In a first example, the predetermined schema is an XSD schema and the data structure is an XML file.

[0085] In a second example, the predetermined schema is a JSON schema and the data structure is a JSON file.< / extensions>

Claims

1. Method (100) for converting a first data file, referred to as the conversion origin file, into a second data file, referred to as the conversion destination file, in order to intercept communications, the conversion origin file having a predetermined proprietary format specific to any one of a plurality of data sources, said sources being sensors suitable for intercepting said data in a telecommunication network, the conversion destination file having a predetermined generic format independent of any one of the plurality of data sources, the predetermined generic format being associated with a predetermined schema which describes the data structure of the predetermined generic format and comprises a plurality of fields, including fields identifying said sources, the method comprising: - a step (110) of obtaining match relationships between the format of the conversion origin file and the format of the conversion destination file, - a step (120) of forming the conversion destination file on the basis of the predetermined schema and the data in the conversion origin file that are involved in the match relationships, - a step (130) of detecting whether there are any remaining data in the conversion origin file that are not involved in the match relationships, referred to as non-matching data, and in this case a step (140) of extending the format of the conversion destination file to include all or some of the non-matching data, wherein the data in the conversion origin file relate to at least one discursive electromagnetic signal which indicates the implementation of at least one communication service in at least one telecommunication network via at least one communication medium and comprise information resulting from the interception of telecommunications, said data comprising identifiers of the service and / or of the communication medium, at least one start time of the communication service, and one or more identifiers of items of equipment from which said signal originates and which receive said signal, including one or more items of user equipment and / or one or more nodes of the telecommunication network that are involved in the implementation of the discursive communication service, and locations of said items of equipment, wherein the forming step (120) comprises a step (121) of identifying, on the basis of the conversion origin file, at least one unique path taken by a communication service which has transited via the node or nodes of the telecommunication network that are involved in the implementation of the discursive communication service, wherein the predetermined schema is arranged to separate the node or nodes of the telecommunication network that are involved in the implementation of the discursive communication service and the unique path or paths taken by the communication service which has transited via the node or nodes of the telecommunication network.

2. Method (100) according to claim 1, wherein the predetermined schema comprises a plurality of fields capable of describing the largest number of data associated with the plurality of data sources.

3. Method (100) according to any one of claims 1 to 2, wherein the predetermined schema is an XSD schema and the step (120) of forming the conversion destination file comprises forming a file in XML format.

4. Method (100) according to any one of claims 1 to 3, wherein the predetermined schema is a JSON schema and the step (120) of forming the conversion destination file comprises forming a file in JSON format.

5. Recording medium recording an instruction program executable by a processor in order to execute all the steps of the method according to any one of claims 1 to 4.

6. Data structure with a predetermined generic format independent of any data source for storing data, the data structure being obtained by a conversion method (100) according to any one of claims 1 to 4 on the basis of a data file and a predetermined schema which defines the structure of the predetermined generic format, said structure comprising a first data structure part and a second data structure part which are separate from each other, the first data structure part comprising a list of the node or nodes of the telecommunication network that are involved in the implementation of the discursive communication service, the second data structure part comprising a list of the unique path or paths taken by a communication service which has transited via the node or nodes of the telecommunication network that are involved in the implementation of the discursive communication service.

7. Data structure according to claim 6, wherein the predetermined schema is an XSD schema and the data structure is an XML file.

8. Data structure according to claim 6, wherein the predetermined schema is a JSON schema and the data structure is a JSON file.