How the use of data for AI training is regulated in Georgia
Training artificial intelligence systems requires vast volumes of records, texts, images and other data that are often collected about other people. This raises a fundamental legal question: may such data be used to teach a model. In Georgian law this question is answered by general norms: the Georgian Law on Personal Data Protection contains no article specifically devoted to artificial intelligence. This must be said plainly: using training data is not in conflict with the law — it is simply subject to the same general regime that governs any other processing: the processing principles, the grounds, the rules on special-category data and the design requirements.
This means that when collecting training data and feeding them to a model, an organisation must answer the same questions that accompany any other processing: what the purpose of processing is, on which ground it takes place, how compatible that purpose is with the original purpose for which the data were collected, and by which measures the risk to the rights of subjects is reduced. The Georgian law follows the European approach, so international practice remains a comparative framework, while the decisive argument is always the Georgian norm.
The compatibility test for further processing
Using data for artificial intelligence training is almost always further processing: the data were collected for one purpose and are used for another — a training — purpose. Under Article 4 of the law, further processing of data for a purpose incompatible with the original purpose is inadmissible. To determine compatibility, the second paragraph of Article 4 establishes six factors that form the basis for assessing a training project.
First: whether a connection exists between the original purpose of collection and the training purpose. Second: the character of the relationship between the controller and the data subject at the time of collection. Third: whether the subject has a reasonable expectation of such further processing of data about him or her. Fourth: whether special-category data are being processed. Fifth: the possible consequences for the subject that may accompany this processing. Sixth: the existence of technical and organisational security measures.
In practice, if data collected from a consumer for the provision of a service unexpectedly become part of a training corpus, the reasonable-expectation factor is the weakest link: the subject could not anticipate such use, and this points towards a conclusion of incompatibility. Documentation of training projects should therefore begin with an analysis against these six factors.
Grounds of processing for training purposes
Article 5 of the law regulates the grounds on which training processing may rest. The data subject's consent to processing for one or several specific purposes is the most reliable ground, provided it satisfies the requirements for consent. Processing is also admissible where it is necessary for the performance of an obligation under a contract with the subject or for entering into a contract at the subject's request; where it is provided for by law; where it is necessary for the performance of duties imposed by legislation; where the data are publicly available under the law or the subject made them publicly available; and in cases of vital interests, significant public interest, tasks in the sphere of public interest, legitimate interests, and the handling of the subject's application.
The legitimate-interests ground is often used in technology projects, but it has a clear boundary: it does not operate where the interests of protecting the rights of the data subject, including a minor, prevail. The second paragraph of the same article distributes responsibility: the duty to substantiate the legal ground of processing rests on the controller. The documentation of a training project is precisely the framework for that substantiation.
Special-category data in training corpora
Special-category data — those connected with health, belief, political views and similar spheres — require special protection. Under the first paragraph of Article 6 of the law, processing of such data is admissible only where the guarantees of protection of the subject's rights and interests provided for by law are ensured by the controller and one of the listed grounds exists. Among those grounds the first is the subject's written consent for one or several specific purposes.
Particularly relevant for training purposes is the ground concerning processing necessary for archiving in the public interest, for scientific or historical research or for statistical purposes in accordance with the law. This ground operates where the law provides for appropriate and specific measures to protect the rights and interests of subjects, and it is not applied where a special law restricts the processing of such data under additional conditions. There is also a ground for ensuring information security and cybersecurity. The third paragraph of Article 6 repeats the same rule: the substantiation of the ground rests on the controller.
By-design protection and minimisation in training systems
Article 26 of the law requires embedding protection when developing new technologies, and an artificial intelligence system is one of the clearest examples of this requirement. Under its first paragraph, taking into account new technologies, the costs of implementation, the nature, scale, context and purposes of processing, as well as the expected risks to the rights of subjects, the controller must adopt appropriate technical and organisational measures, including pseudonymisation, both when determining the means and directly in the process.
The second paragraph governs the priority of more extensive masking of data: when determining the quantity of data, the scale of processing, retention periods and access, it must be ensured that only the volume necessary for the specific purpose is processed automatically. For training corpora this means in practice: depending on the model's task, preventing unnecessary fields, identifiers and special-category data from entering the corpus, and incorporating pseudonymisation at an early stage of the data flow.
What an organisation should do before a training project
The correct sequence is as follows. First: the original purpose of collection of each data flow is documented. Second: the training purpose is analysed against the six factors of the second paragraph of Article 4. Third: a ground is selected from Article 5, and for special-category data the corresponding ground from Article 6, including written consent. Fourth: minimisation and pseudonymisation are designed in accordance with Article 26. Fifth: the substantiation of the ground and the analysis are documented so that the organisation can present them at the appropriate moment.
The Legal.ge team will help you with the legal assessment of training data: from describing data flows and analysing compatibility to selecting grounds and documenting design measures, so that your artificial intelligence project rests on the precise requirements of Georgian legislation.
