Showing posts with label technology assisted review. Show all posts
Showing posts with label technology assisted review. Show all posts

Thursday, March 27, 2014

Jason Atchley : eDiscovery : Automated vs. Linear Review, Which One Makes Sense?

jason atchley

Automated vs. Linear Review, Which One Makes Sense?

Sorting out the facts and dispelling the myths of predictive coding schemes.
, Law Technology News
    |0 Comments

Laptop and computer files in folders on white isolated background. 3d
E-discovery has become a tricky and complex process for attorneys.  While automated review can be a good option, Huron Legal Managing Director Nathalie Hofman, who advises law departments and firms on document review, says there are still times when linear review is the better choice.  In a recent Eye on Discovery piece on the company’s website, Hofman did an internal Q & A with Robert G. Kidwell, member of Mintz, Levin, Cohn, Ferris, Glovsky and Popeo. The interview offers legal professionals some guidelines on when to use, and not use, automated review.   
Kidwell advises against using automated review in instances where there are complex issues or similarity between identically responsive documents. On the hand, he says it’s very helpful in a situation where you have a very broad subpoena or other discovery obligation and “it’s basically an ‘all business documents’ request.”
He discusses an instance in which the issues separating a responsive document from a nonresponsive one were not that clear-cut to allow pure semantic logic to distinguish between the two.  In the end he says human input was necessary to determine what the dispute was about, the types of documents that might be relevant, etc., adding that this proved to be faster and more effective than the automated method which he also employed.
Another tip, he says don’t discount the usefulness of search terms in automated review. 
At the end of the day, Kidwell says both forms of document review remain relevant; although powerful, automated review is not a panacea.
Sherry Karabin is a freelance writer and reporter based in New York City. Email: sherry.karabin@yahoo.com.


Read more: http://www.lawtechnologynews.com/id=1395850114452/Automated-vs.-Linear-Review%2C-Which-One-Makes-Sense%3F#ixzz2xAaNKRGR



Monday, March 24, 2014

Jason Atchley : Kroll Ontrack : Step Out of the Predictive Coding Rain!

jason atchley

PREDICTIVE CODING SEMANTICS: STEP OUT OF THE RAIN!

Wednesday, February 19, 2014

Jason Atchley : eDiscovery : Predictive Coding is SOOOO Yesterday!

jason atchley

Predictive Coding Is So Yesterday

Computational linguistics and data mining are among the tools that will drive e-discovery in coming years.
Law Technology News
    |0 Comments

man reading computer printout.
Editor's Note: This article was chosen in a blind competition by the Arizona State University-Arkfeld E-Discovery and Digital Evidence Conference. The three winners have been invited to present their papers during the conference, which will be held March 12-14 at ASU's Sandra Day O'Connor College of Law, in Tempe, Ariz. See also, "Vendor Voice: Yes, Counselor, There Will Be Math."
Today, much of the electronic data discovery industry is racing to improve predictive coding, which is but one approach to technology-assisted review—and one with inherent limitations.  Instead of refining predictive coding, tomorrow’s innovative EDD technology game changers will employ computational linguistics, data mining, language translation, corpus-based content analysis and case specific information supplied in the form natural language inquiries.  
This is not to say others haven’t attempted to apply these technologies to EDD.  However, significant advances in quality, and cost-reducing innovations, will be driven by the integration of techniques from these disciplines.
Predictive coding depends critically on the creation of “training sets,” created by one or more human reviewers through manual review. The quality of these training sets entirely determines the recall and precision achieved by the technology, because, to date, predictive coding tools apply information from training sets but do not correct reviewer errors within training sets. While current offerings use a variety of methods for selecting electronically stored information to be reviewed, none of the methods can actually assist the user in making correct markings. The lack of analysis taking place in the front-end of today’s predictive coding offerings place an upper limit on predictive coding effectiveness.
Once predictive coding applies a training set to a population of unreviewed ESI, a set of human reviewers must once again review a selected set of ESI marked by the predictive coding technology. Once again the human reviewer markings of ESI critically determines the accuracy and completeness of the next round of predictive coding markings of the unreviewed ESI.  In reality, a human reviewer can inconsistently mark ESI during creation of a training set or during subsequent review of predictive coding markings without any feedback or error checking.
Perhaps the most powerful claims of predictive coding are also the most damning. The fact that it  is equally as accurate and complete as human review is best evaluated the way all technology is evaluated. Technology makes our lives better—we travel faster, hear better, see farther, lift more weight, and drill smaller holes, than we possibly could without it. The reality that predictive coding enables us to review more ESI than if done entirely by human reviewers certainly is true, yet that claim stems from large storage capabilities and fast processor speeds executing the predictive coding tools, not the predictive coding technology itself.
Predictive coding  brings limitations with its advantages. The dependence upon the accuracy of human review, review that takes place without feedback or error checking, will limit the recall and precision of predictive coding until some type of pre-processing is done to relate ESI content and thus perform some type of error checking.  Without ESI content analysis and relationship identification, human review errors propagate into the technology, especially if these errors occur consistently.

The Future of Predictive Coding

Technology-assisted review tools of the future will analyze ESI using computational linguistics without any user input to analyze ESI content, far beyond keywords, key phrases, or training sets of ESI documents subject to human error. Content analysis will allow powerful categorization of ESI based on data mining and language translation techniques.
Once categorized, users can review categories of ESI rather than any set number of ESI items (as is required by predictive coding). Human review will take place under error checking and marking consistency feedback made possible through information theory measures drawn from categories of semantic meaning. Such meaning will be determined by the content of ESI populations and not any error-prone training of predictive coding.
Because these methods analyze ESI based on each the semantic content of ESI data sets, such methods will be adaptable to a wide range of ESI content and in fact be ESI content driven. In other words, there will be no training sets—or user defined categories—for the algorithms to learn about or compare.
The categories derived from the semantic meaning of ESI within data sets will support fast corpus-based analysis by human reviewers. As human reviewers mark semantic categories, rather than individual ESI, the systems of tomorrow will concurrently compare reviewer actions and markings to existing categories and markings. Comparison will provide feedback to improve the review process in real-time, not overnight through a learning process.
This comparison will provide feedback to the human reviewer to assist the reviewer in taking correct actions and making consistent markings. Such on-the-fly feedback and consistency checking will elevate human reviewers to more powerful reviewers—increasing the accuracy, speed, and consistency of review. Rather than a predictive coding tool working to understand user actions and making the best of user errors and inconsistencies, the technology will create a super-user, able to produce better results faster and cheaper. In addition, if new issues arise, the super-user need only return to the analyzed and categorized ESI to investigate and locate pertinent ESI—no new training set need be created and there is no need for yet another cycle of “review-train-revise-train…” to suffer through, wait for, and pay for.
The quality future technologies will be uniquely based on the ESI, the analysis and categorization algorithms, and the human reviewer—transformed into a super-user—will benefit from error checking and consistency measures. The integration of these multiple fields will bring new tools to TAR just as the introduction of new technologies expanded power, speed, and other abilities of humans. In fact, these innovators not bound by years of investment in predictive coding are bringing new technologies to the market today. These tools will make predictive coding the analog technology of yesterday.
Joel Henry is an attorney professor of computer science and IT legal advisor at the University of Montana, based in Missoula.


Read more: http://www.lawtechnologynews.com/id=1202643337112/Predictive-Coding-Is-So-Yesterday#ixzz2tmZskGLo



Monday, February 17, 2014

Jason Atchley : eDiscovery : Statistics, Rule 26(g) and TAR

jason atchley

Vendor Voice: Statistics, Rule 26(g) and Getting Stuck in TAR

The TAR process must be implemented with a consistent eye toward certification requirements.
, Law Technology News
    |0 Comments

Anyone who has ever tried to use a technology-assisted review or predictive coding tool usually starts by talking to a vendor—or a handful of vendors—who immediately suggest these tools are exceedingly simple to use and speed up the time and lower the costs of litigation. While no doubt true, attorneys in federal court are held to a standard of “reasonable inquiry” as dictated by Rule 26(g) of the Federal Rules of Civil Procedure. If attorneys do not keep a mindful eye on the process, the easy button of TAR can raise unintended Rule 26(g) challenges by the opposing party or unilaterally by the Judge in the case. 
The Implications of Rule 26(g) on the Use of Technology-Assisted Review was recently published in the Federal Courts Law ReviewThe article analyzes five phases of the TAR process that, if not fully  considered and properly executed can engender Rule 26(g) arguments. These stages are collection, disclosure, training, stabilization, and validation. For example, during the collection phase, attorneys seldom consider the impact of the richness of the collection upon the reasonableness of the inquiry under Rules 26(g).  Since the advent of the recent federal rules and the warnings of Zubulake V (Zubulake v. UBS Warburg 229 F.R.D. 422 (S.D.N.Y. 2004), attorneys have been fearful of sanctions for not preserving and collecting all relevant electronically stored information. The knee jerk reaction has been to preserve and collect broadly, and then throw more data into a review tool than is even remotely tied to a case. This is compounded by the fact that requesting parties, recognizing the relative ease of searching ESI (as compared with hard copy documents) tend to make overly broad document production requests. 
As a practical matter, this can make it more difficult to implement a technology-assisted review, which depends on the development of a language model to distinguish between relevant and non-relevant documents. This is generally accomplished by the algorithmic analysis of language patterns in documents which are coded responsive versus documents which are coded not responsive. If relatively few of the documents in a collection are responsive, it becomes a challenge finding enough documents to develop the model. If you use a seeding approach of finding and picking exemplar documents, much like we use key words, you run the risk of not finding the documents that, while perhaps relevant and even important to the case, are not sufficiently like the seed documents to be uncovered by the tool.
If you use a machine assisted or random approach, it may be necessary to code a significant number of documents to develop the model. For example, if only 1 percent of the collection is relevant, a completely random selection of documents may require a review of 20,000 to 50,000 documents. While this will typically be a very small fraction of the entire collection, it can be difficult for senior attorneys to devote the necessary time to the review.
This situation also implicates the Rule 26(g) certification. TAR productions are often validated to a confidence level of 95 precent and a confidence interval of ±2 percent by reviewing just under 2,400 documents. For a reasonable collection in which 10 percent of the documents may be relevant, the actual confidence interval would be closer to ±1.2 percent, or just 12 percent of the anticipated value.
However, in a poor collection in which only one percent of the documents are expected to be relevant, the confidence interval, while only ±0.4 percent, would actually equate to 40 percent of the estimated value. The article discusses the implications of this situation on the Rule 26(g) certification, because the producing party is uniquely situated to manage the results of the TAR process. 
The article also addresses other situations, such as the challenges implicit in effecting cooperation and transparency when lawyers are typically accustomed to sharing as little information as necessary with opposing counsel. In cases such as In Re Actos and Global Aerospace v. Landow Aviation, L.P. dba Dulles Jet Center, et al, the parties agreed to share documents coded as non-responsive which were fed into the TAR tool, in order to gain an agreement from opposing counsel and to reduce the risks of a challenge to  the training process.
 The concept behind sharing this information is similar to the idea behind sharing key words with opposing counsel. Any agreement reduces the risk of being challenged for not having undertaking a “reasonable inquiry.” Cases spawned by 75 years of manual review do not require this level of transparency and cooperation, and courts have been slow to move in that direction. See In Re Biomet. Nevertheless, deficiencies in training the TAR tool, can only be discovered by an adversary who has not been the beneficiary of transparency and cooperation after the time and expense of training have been incurred by the defendants. Any deficiencies in the production may be viewed negatively under Rule 26(g) if the court sees transparency as a way to  reduce the cost of litigation and improve the value of discovery. This is especially true if the producing party opposed transparency during the course of the litigation. The article explains how transparency can serve as  an insurance policy against a Rule 26(g) challenge.
The key takeaway: Rule 26(g) implicates every aspect of the TAR process, and the process must be implemented with a consistent eye toward certification requirements. Equally as important is the notion that linear review (which is typically conducted on electronic data) may well become subject to the same types of considerations as attorneys attempt to impose validation requirements on modern document productions. As Rule 26(g) moves to the forefront, lawyers who do not fully appreciate the background sampling and statistics risk finding themselves stuck in “tar” of the reasonable inquiry standards of Rule 26(g)
Attorney Karl Schieneman is president of Review Less (kas@reviewless.com); Thomas Gricks III is head of predictive coding and a shareholder at Schnader Harrison Segal and Lewis (tgricks@schnader.com). Both are based in Pittsburgh, Pa., and participated in the Global Aerospace case.


Read more: http://www.lawtechnologynews.com/id=1202642979755/Vendor-Voice%3A-Statistics%2C-Rule-26%28g%29-and-Getting-Stuck-in-TAR#ixzz2tb8nlGSB