U.S. law firms may be worried about the security risks of sharing confidential information online, but a new surveyby LexisNexis' legal and professional division reveals that they are not doing much about it.
Unencrypted email remains by far the most prominent way that law firms share privileged communications with their clients, with 89 percent of respondents reporting that it is the firm's primary method of distributing information.
In March, the company canvassed about 300 legal professionals in 40 states across 15 different practice areas. Results show that although respondents were aware of the risks, and wary of them, the most common method of securing documents and protecting privilege was the use of a confidentiality statement at the bottom of an email, with 77 percent of firms reporting this was their primary line of defense.
“There’s clearly a disconnect between expressed security concerns and measures law firms employ to protect their clients and themselves,” said Christopher Anderson, a senior product manager at LexisNexis, in a statement. “Relying on a mere statement of confidentiality when sharing privileged communications by email is a weak measure—and further it might protect the law firm but affords very little protection for the client,” he said.
A minority of law firms go a step further to protect their information, with 22 percent saying they use email encryption,14 percent using a password to protect documents and 13 percent employing a secure file-sharing site. At the reverse end of the spectrum, 4 percent of respondents said they take no measures at all to protect private information. “Law firms need to perform their due diligence, stay abreast of technology and ultimately protect their clients’ interest online just as they do in providing legal counsel,” said Anderson.
Attorney Marlisse Silver Sweeney is a freelance writer based in Vancouver. Twitter: @MarlisseSS.
If you worry about where your data is today, do you have a plan for when you wear your data to and from work?
The dangers of wearable devices are not imminent, and many technologies (e.g., Google Glass) are still in beta, according to a January Forrester Research Inc. report called “The Enterprise Wearable Journey,” by J.P. Gownder, vice president and principal analyst focusing on infrastructure and operations professionals at Forrester. The wearable industry will blossom over the next decade, the report says.
However, law firms can start taking steps now to avoid leaking client data through wearable devices, according to industry analysts.
“We should start today familiarizing ourselves with what is different about the wearables and the [Internet of Things] from traditional security—forearmed is forewarned,” said Earl Perkins, research vice president at Gartner Inc., in an LTN interview. “It’s a new, complex world.”
Stamford, Conn.-based research and advisory firm Gartner anticipates wearable electronic revenue worldwide will more than triple in the years leading up to 2016, from $1.6 billion to $5 billion.
Although there has not been a notable case to date where a hacker used wearables as an entry point to infiltrate an internal network, according to Gartner’s Perkins, there is a possibility such an incident could occur.
Most data collected from wearable technology today is not encrypted and transmitted with “minimal protection,” Perkins said.
To avoid an information hazard, firms using wearable devices should ensure data are encrypted, says Sharon Nelson, president of Sensei Enterprises Inc., an IT, digital forensics and information security firm located in Fairfax, Va.
Most wearable devices use Bluetooth or Wi-Fi, Nelson said. Wi-Fi must be encrypted with WEP2 (Wired Equivalent Privacy), and Bluetooth should be disabled on law firm devices when they are not in use, she says. Bluesnarfing, or stealing information though a Bluetooth connection, will likely become more common as wearable devices increase in popularity, according to Nelson.
“We have already heard anecdotally that some law firms are banning Google Glass,” Nelson said.
Editor's Note: This article was chosen in a blind competition by the Arizona State University-Arkfeld E-Discovery and Digital Evidence Conference. The three winners have been invited to present their papers during the conference, which will be held March 12-14 at ASU's Sandra Day O'Connor College of Law, in Tempe, Ariz. See also, "Vendor Voice: Yes, Counselor, There Will Be Math."
Today, much of the electronic data discovery industry is racing to improve predictive coding, which is but one approach to technology-assisted review—and one with inherent limitations. Instead of refining predictive coding, tomorrow’s innovative EDD technology game changers will employ computational linguistics, data mining, language translation, corpus-based content analysis and case specific information supplied in the form natural language inquiries.
This is not to say others haven’t attempted to apply these technologies to EDD. However, significant advances in quality, and cost-reducing innovations, will be driven by the integration of techniques from these disciplines.
Predictive coding depends critically on the creation of “training sets,” created by one or more human reviewers through manual review. The quality of these training sets entirely determines the recall and precision achieved by the technology, because, to date, predictive coding tools apply information from training sets but do not correct reviewer errors within training sets. While current offerings use a variety of methods for selecting electronically stored information to be reviewed, none of the methods can actually assist the user in making correct markings. The lack of analysis taking place in the front-end of today’s predictive coding offerings place an upper limit on predictive coding effectiveness.
Once predictive coding applies a training set to a population of unreviewed ESI, a set of human reviewers must once again review a selected set of ESI marked by the predictive coding technology. Once again the human reviewer markings of ESI critically determines the accuracy and completeness of the next round of predictive coding markings of the unreviewed ESI. In reality, a human reviewer can inconsistently mark ESI during creation of a training set or during subsequent review of predictive coding markings without any feedback or error checking.
Perhaps the most powerful claims of predictive coding are also the most damning. The fact that it is equally as accurate and complete as human review is best evaluated the way all technology is evaluated. Technology makes our lives better—we travel faster, hear better, see farther, lift more weight, and drill smaller holes, than we possibly could without it. The reality that predictive coding enables us to review more ESI than if done entirely by human reviewers certainly is true, yet that claim stems from large storage capabilities and fast processor speeds executing the predictive coding tools, not the predictive coding technology itself.
Predictive coding brings limitations with its advantages. The dependence upon the accuracy of human review, review that takes place without feedback or error checking, will limit the recall and precision of predictive coding until some type of pre-processing is done to relate ESI content and thus perform some type of error checking. Without ESI content analysis and relationship identification, human review errors propagate into the technology, especially if these errors occur consistently.
The Future of Predictive Coding
Technology-assisted review tools of the future will analyze ESI using computational linguistics without any user input to analyze ESI content, far beyond keywords, key phrases, or training sets of ESI documents subject to human error. Content analysis will allow powerful categorization of ESI based on data mining and language translation techniques.
Once categorized, users can review categories of ESI rather than any set number of ESI items (as is required by predictive coding). Human review will take place under error checking and marking consistency feedback made possible through information theory measures drawn from categories of semantic meaning. Such meaning will be determined by the content of ESI populations and not any error-prone training of predictive coding.
Because these methods analyze ESI based on each the semantic content of ESI data sets, such methods will be adaptable to a wide range of ESI content and in fact be ESI content driven. In other words, there will be no training sets—or user defined categories—for the algorithms to learn about or compare.
The categories derived from the semantic meaning of ESI within data sets will support fast corpus-based analysis by human reviewers. As human reviewers mark semantic categories, rather than individual ESI, the systems of tomorrow will concurrently compare reviewer actions and markings to existing categories and markings. Comparison will provide feedback to improve the review process in real-time, not overnight through a learning process.
This comparison will provide feedback to the human reviewer to assist the reviewer in taking correct actions and making consistent markings. Such on-the-fly feedback and consistency checking will elevate human reviewers to more powerful reviewers—increasing the accuracy, speed, and consistency of review. Rather than a predictive coding tool working to understand user actions and making the best of user errors and inconsistencies, the technology will create a super-user, able to produce better results faster and cheaper. In addition, if new issues arise, the super-user need only return to the analyzed and categorized ESI to investigate and locate pertinent ESI—no new training set need be created and there is no need for yet another cycle of “review-train-revise-train…” to suffer through, wait for, and pay for.
The quality future technologies will be uniquely based on the ESI, the analysis and categorization algorithms, and the human reviewer—transformed into a super-user—will benefit from error checking and consistency measures. The integration of these multiple fields will bring new tools to TAR just as the introduction of new technologies expanded power, speed, and other abilities of humans. In fact, these innovators not bound by years of investment in predictive coding are bringing new technologies to the market today. These tools will make predictive coding the analog technology of yesterday.
Joel Henryis an attorney professor of computer science and IT legal advisor at the University of Montana, based in Missoula.
Anyone who has ever tried to use a technology-assisted review or predictive coding tool usually starts by talking to a vendor—or a handful of vendors—who immediately suggest these tools are exceedingly simple to use and speed up the time and lower the costs of litigation. While no doubt true, attorneys in federal court are held to a standard of “reasonable inquiry” as dictated by Rule 26(g) of the Federal Rules of Civil Procedure. If attorneys do not keep a mindful eye on the process, the easy button of TAR can raise unintended Rule 26(g) challenges by the opposing party or unilaterally by the Judge in the case.
The Implications of Rule 26(g) on the Use of Technology-Assisted Review was recently published in the Federal Courts Law Review. The article analyzes five phases of the TAR process that, if not fully considered and properly executed can engender Rule 26(g) arguments. These stages are collection, disclosure, training, stabilization, and validation. For example, during the collection phase, attorneys seldom consider the impact of the richness of the collection upon the reasonableness of the inquiry under Rules 26(g). Since the advent of the recent federal rules and the warnings of Zubulake V (Zubulake v. UBS Warburg 229 F.R.D. 422 (S.D.N.Y. 2004), attorneys have been fearful of sanctions for not preserving and collecting all relevant electronically stored information. The knee jerk reaction has been to preserve and collect broadly, and then throw more data into a review tool than is even remotely tied to a case. This is compounded by the fact that requesting parties, recognizing the relative ease of searching ESI (as compared with hard copy documents) tend to make overly broad document production requests.
As a practical matter, this can make it more difficult to implement a technology-assisted review, which depends on the development of a language model to distinguish between relevant and non-relevant documents. This is generally accomplished by the algorithmic analysis of language patterns in documents which are coded responsive versus documents which are coded not responsive. If relatively few of the documents in a collection are responsive, it becomes a challenge finding enough documents to develop the model. If you use a seeding approach of finding and picking exemplar documents, much like we use key words, you run the risk of not finding the documents that, while perhaps relevant and even important to the case, are not sufficiently like the seed documents to be uncovered by the tool.
If you use a machine assisted or random approach, it may be necessary to code a significant number of documents to develop the model. For example, if only 1 percent of the collection is relevant, a completely random selection of documents may require a review of 20,000 to 50,000 documents. While this will typically be a very small fraction of the entire collection, it can be difficult for senior attorneys to devote the necessary time to the review.
This situation also implicates the Rule 26(g) certification. TAR productions are often validated to a confidence level of 95 precent and a confidence interval of ±2 percent by reviewing just under 2,400 documents. For a reasonable collection in which 10 percent of the documents may be relevant, the actual confidence interval would be closer to ±1.2 percent, or just 12 percent of the anticipated value.
However, in a poor collection in which only one percent of the documents are expected to be relevant, the confidence interval, while only ±0.4 percent, would actually equate to 40 percent of the estimated value. The article discusses the implications of this situation on the Rule 26(g) certification, because the producing party is uniquely situated to manage the results of the TAR process.
The article also addresses other situations, such as the challenges implicit in effecting cooperation and transparency when lawyers are typically accustomed to sharing as little information as necessary with opposing counsel. In cases such as In Re Actos and Global Aerospace v. Landow Aviation, L.P. dba Dulles Jet Center, et al, the parties agreed to share documents coded as non-responsive which were fed into the TAR tool, in order to gain an agreement from opposing counsel and to reduce the risks of a challenge to the training process.
The concept behind sharing this information is similar to the idea behind sharing key words with opposing counsel. Any agreement reduces the risk of being challenged for not having undertaking a “reasonable inquiry.” Cases spawned by 75 years of manual review do not require this level of transparency and cooperation, and courts have been slow to move in that direction. SeeIn Re Biomet. Nevertheless, deficiencies in training the TAR tool, can only be discovered by an adversary who has not been the beneficiary of transparency and cooperation after the time and expense of training have been incurred by the defendants. Any deficiencies in the production may be viewed negatively under Rule 26(g) if the court sees transparency as a way to reduce the cost of litigation and improve the value of discovery. This is especially true if the producing party opposed transparency during the course of the litigation. The article explains how transparency can serve as an insurance policy against a Rule 26(g) challenge.
The key takeaway: Rule 26(g) implicates every aspect of the TAR process, and the process must be implemented with a consistent eye toward certification requirements. Equally as important is the notion that linear review (which is typically conducted on electronic data) may well become subject to the same types of considerations as attorneys attempt to impose validation requirements on modern document productions. As Rule 26(g) moves to the forefront, lawyers who do not fully appreciate the background sampling and statistics risk finding themselves stuck in “tar” of the reasonable inquiry standards of Rule 26(g)
Attorney Karl Schieneman is president of Review Less (kas@reviewless.com); Thomas Gricks III is head of predictive coding and a shareholder at Schnader Harrison Segal and Lewis (tgricks@schnader.com). Both are based in Pittsburgh, Pa., and participated in the Global Aerospace case.
You’re heading out for a meeting with a client when suddenly you realize you can’t find an important document. Did it get deleted or simply misplaced? Scenarios like this are one example why Baker & Hostetler attorneys Judy Selby and James Sherer are urging firms and others in the legal profession to get their data houses in order. In a blog posted in Information Security on the firm’s website they discuss “Information Governance.”
Selby and Sherer argue data security concerns, privacy, compliance and e-discovery costs are just some of the reasons that sound policies to efficiently manage information must be in place. Their key points:
Policy must be consistent with “enterprise-wide strategic and business goals,” they say. It should include “all relevant stakeholders and take into account the enterprise’s organization and culture, legal/regulatory concerns, business operations and technology.”
Special data challenges like the retention of personal health information or the management of streaming social media data must be addressed.
Most data likely has no business value. Implement a defensible deletion plan guided by considerations, such as the effect of legal holds, regulatory and compliance requirements etc.
Include guidelines for management of retained information and eliminate redundancies, creating classification and organizational systems so things can be retrieved quickly.